Trial it, don't move.
On the only two independent same-harness boards, K3, Opus 4.8 and Fable 5 are inside each other's noise. Terminal-Bench 2.1, one evaluator, one scaffold: K3 85.0% · Fable 5 84.6% · Opus 4.8 84.6% — K3 ahead of both by one task instance. DeepSWE, one scaffold: Fable 5 70%±4 · K3 69%±5 — overlapping intervals, at $21.63 vs $4.65 per task. But that cost gap is an API-list-price gap. The practitioners who ran both on subscriptions report it vanishing — “K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan” — and it is now measured — by two people, and they are not the same person: u/djdante reports that at $20 a side, K3 consumes 4× the weekly quota of Fable 5 High for the same task (⚠ his own words: he ran those tests “for my YouTube channel”, and no artefact, harness or task set is published — disclosed because it costs us, not because it helps) 🕐 and read the date: that was measured 2026-07-15, against a $20 Claude Pro plan that still included Fable 5. Pro stopped including it on 2026-07-20. The measurement stands as made; the configuration does not. To run it as of 2026-07-20 you would need a $100 Max plan on the Claude side; u/ahnyudingslover, replying directly to him on a production codebase, adds that output quality “was close, which was shocking” (a direct reply opening “I can attest to this as well” — corroboration in-thread, not a second independent measurement, and also a $20-tier report carrying the same pre-07-20 dating caveat). So if you pay by plan rather than by token, the cheap-K3 case is not merely unproven — it points the other way. (It reverses again at $100, where the same complaint becomes “very generous”. Tier matters.) So the thing that actually separates these models is not the score. It is what you can buy: K3's weights aren't out, its licence is unannounced, and Moonshot paused new subscriptions on 07-19 — while on the Claude side, from 2026-07-20, Fable 5 is a permanent part of Max and Team Premium only, and there at 50% of each plan's usage limits. Max is an individual plan — it starts at $100/month and runs to $200. Claude Pro ($20) and Team Standard lost their included Fable 5 on that date: they keep it only through usage credits billed at API rates ($10/M in, $50/M out) after a one-time $100 credit. And the cap compounds — the 50% boost to Claude Code weekly limits ended the same day, so the half-slice is measured against a baseline the-decoder puts ~33% lower. “Just use Fable 5” is still not a $20-plan answer — it is a $100-plan answer, at half of a limit that just shrank. Both sides have an access problem, and neither has a settled quality gap. That is a timing answer. Run K3 as a second seat on the shapes it demonstrably wins. Leave the loop you ship from where it is. ⚠ And read our confidence precisely, because the two halves differ: this card is LOW because the quality and harness questions are wide open — the K3–Fable gap is one instance in 267 on one board and inside the error bars on the other, pointing opposite ways, and no independent test in any language — across every source in this sweep — has run the same tasks in both a Kimi and a Claude-shaped harness. But the subscription-economics finding specifically is close to settled: one quantified $20-vs-$20 ratio plus first-hand corroboration on six platforms in three languages. (That used to read “two quantified measurements”. It was an over-claim: djdante gives a ratio, ahnyudingslover gives only Fable's own consumption and no K3 figure. One ratio, corroborated in-thread.) LOW is a judgement about the verdict, not about that finding. ⚠️ And a warning about that word “settled”: a third verifier found this card falsified on its own publication date — on the access half, the half we were calling near-settled — because Anthropic changed which plans include Fable 5 on 2026-07-20 and we had not read it, in two sources this card already cites. It is corrected above. Treat “near-settled” as a statement about a moving object, and check the plan tables yourself before you spend anything.
The decision
CHOOSE CLAUDE IF →
- Your work is greenfield frontend, visual or 3D — Frontend Code Arena #1 at 1679, “surpassing Claude Fable 5” and #1 in 6 of 7 domains (Arena's own post); the relayed leaderboard puts Fable 5 at 1631 and Opus 4.8 (Thinking) at 1562. Corroborated by testers otherwise hostile to the model
- Cost per finished task binds, you bill at API rates, and you can get a seat — same-harness DeepSWE: $4.65 vs Opus 4.8's $13.22 and Fable 5's $21.63, on fewer output tokens than either (81k vs 135k vs 119k). This advantage is an API-list-price advantage and does not survive onto a subscription — see the stay-put column
- You want frontier-ish coding on a cheap plan — because from 2026-07-20 the cheapest Claude plan that includes Fable 5 is Max, at $100/month (Max 5× $100 · Max 20× $200 · Team Premium $100/seat), and even there it is capped at 50% of that plan's usage limits, against a weekly baseline that shrank the same day. On Pro ($20) and Team Standard, Fable 5 is now credits-only at API rates — $10/M in, $50/M out — after a one-time $100 credit. Opus 4.8 needs Pro or above. K3 at $3 / $0.30 cached / $15 per MTok is a third of Fable 5's $10 / $1 / $50
- You'll run it in Kimi Code, or a harness you have verified echoes thinking history back — the vendor's own words are that quality “may become highly unstable” otherwise
- You need security analysis a Claude model refuses — a private, un-gameable cyber eval rated K3 “best price/recall” (best value for recall, not best recall) and recommended it for continuous analysis, with Fable 5 at a 100% refusal rate
CHOOSE KIMI K3 IF →
- You want the gap to be real before you move — K3's edge over Fable 5 is one task instance on one board and a reversal inside overlapping error bars on the other. Nothing here clears noise in either direction
- Your work is maintenance on an existing large repo — an independent repair harness put K3 last of 7 (79%), with both Claude models ahead of it
- Your harness is Claude-shaped and you can't verify the thinking-history echo — one live session-wedging bug that names Kimi — but the PR names Moonshot/Kimi and ZAI and carries both provider labels, so it is a two-provider wire-format bug, not K3-specific. Anthropic-style thinking blocks are one of the entry paths it enumerates, not the cause. Hit four times in one day. (This line used to read “three live bugs in three days”. Checked one at a time against the GitHub API on 2026-07-20: the second was closed by its own filer nine minutes after he opened it, and the third is a gateway bug hitting fifteen models with K3 as one list entry. One bug, not three.) The sharpest first-hand report is not on GitHub at all — a B站 user wiring Kimi into Claude Code via
ccswitchwith the model set tok3, asking why k2.6 is what actually gets called. No answer to him appears in the 107 replies we read (pages 1–3 of 372, hot-sorted) — the video's full comment set is login-walled, so we cannot say nobody answered, only that no answer is in what we read - You're on a subscription, not list API — this is the single strongest argument against the whole cheap-K3 case, and it is first-hand: “Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan — i observe no efficiency benefit” (@kunchenguid, on the $40 plan). Corroborated independently: a measured ¥199 Kimi Code plan came out at roughly a Claude Pro Fable allowance (🕐 measured 2026-07-17 — that comparator is gone: Claude Pro stopped including Fable 5 on 2026-07-20), and a 94% cache-hit user still burned the 5-hour window in two tasks. Both vendors meter on a rolling 5-hour window plus weekly caps
- You need the weights, a licence, or a seat — as of 2026-07-20 none of the three existed: no HF repo, licence unannounced, new subscriptions paused
Works with your setup?
| vs Opus 4.8 same harness | vs Fable 5 same harness | Claude-shaped harness | Best model on a cheap plan | Weights / licence | Can you buy a seat? | |
|---|---|---|---|---|---|---|
| Kimi K3 | ⚠ +1 instance / +10 pts | ⚠ tie, ~1/4 cost | ⚠ run, never measured | ⚠ ¥199 for 1M ctx | ✗ neither yet | ✗ paused 07-19 |
| Claude (Opus 4.8 / Fable 5) | peer, 2.8× cost | ⚠ tie, ~4× cost | ✓ it's native | ✗ Fable needs $100+ Max, at 50% limits | ✗ closed, stated | ✓ on sale |
Sentiment — independent voices only
Weighted evidence
- The claim that framed every launch article — K3 “mostly beating Claude Opus 4.8 max… while losing out to Claude Fable 5” — is Simon Willison reading Moonshot's self-reported table, and he says so. Neither independent same-harness board reproduces the Fable loss. simonwillison.net · 2026-07-16 · explicitly “their self-reported benchmarks”
- The strongest same-harness numbers, all three models on mini-SWE-agent, 113 contamination-free tasks: K3 69%±5% at $4.65/task · Fable 5 70%±4% at $21.63 · Opus 4.8 59%±2% at $13.22. K3 ties Fable inside the error bars at ~1/5 the cost, and clears Opus outright DeepSWE / Datacurve · re-fetched 2026-07-20 · “all models run on mini-swe-agent for consistency”
- The Claude-specific footgun, live: sessions permanently wedged, hit “four times in one day on kimi-k3 via OpenRouter” — one entry path being Anthropic-style content thinking-blocks stripped by the wire converter. This is the one genuinely live session-wedging bug we found that names Kimi at all — but it is not K3-specific, and we apply the same discount here that we applied to the gateway bug below: the PR body opens “Strict OpenAI-compatible providers (Moonshot/Kimi, ZAI) reject any request…” and it is labelled
provider/kimiandprovider/zai. It is a two-provider wire-format class bug that K3 users hit, not a defect unique to K3. (This line previously read “genuinely K3-specific” — corrected 2026-07-20 against the PR's own first sentence.) — the card previously claimed three (see the coverage note) hermes-agent #67386 · 2026-07-19 · open · ⚠️ READ THE CAVEATS BEFORE THE BUG: the PR body ends🤖 Generated with Claude Code— this report is AI-authored, and its author is an outside contributor on a fork (author_association: NONE). Repo-labelledprovider/kimi,P2; 6 of its 9 labels are ordinary repo taxonomy, only 3 carry asweeper:prefix, so “bot-applied” describes a third of them, not all - The strongest argument against the cheap-API framing, and it is first-hand: “i don't care what the benchmark numbers say, and what the face value API pricing is, in reality Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan — i observe no efficiency benefit” — on the $40 plan, a few prompts into one 200k-context session took a third of his 5-hour limit. Plus instructions ignored that “were never a problem with gpt 5.5, 5.6, opus, fable and grok 4.5”. List price is not what a subscriber pays. And the same post says this, which we were leaving out: “the pure ‘intelligence’ of K3 does hold up - it understands my intent very well, and can diagnose problems, delegate tasks all fine”. His cost case is untouched by it; quoting only a witness's negatives is how a card starts lying without saying anything false @kunchenguid · 2026-07-17 · builds a rival agent, no stake in either vendor · re-verified verbatim 2026-07-20 — x.com now reads through agent-reach, so you can check this yourself
- The number this card has wanted since it was written — a $20-plan-to-$20-plan measurement, and it now needs a date stamp. “If you have a $20 Kimi plan and a $20 Claude plan - Kimi k3 will consume 4X more of your weekly quota than Fable 5 High for the same task”. 🕐 MEASURED 2026-07-15, AND THE CLAUDE SIDE OF THAT PAIR CHANGED ON 2026-07-20: a $20 Claude Pro plan no longer includes Fable 5 — it is Max ($100–$200) and Team Premium only, at 50% of limits (see the access item in the dropdown). The measurement is real, was valid when made, and is not withdrawn — but it compares against a tier that stopped existing in that shape on this card's publication date, and the same caveat attaches to every $20-tier figure below. Independently, on a production codebase: “3 different exact prompt comparisons between fable5 and Kimi k3… fable5 uses about 30-50% of my 5 hour usage limit on the 20usd plan for both. Output quality was close, which was shocking”. Read those together and they are the whole card: quality close enough to shock a sceptic, subscription economics several times worse. Counterweights, carried: at $100 the same complaint reverses (“very generous”), one user reports Low reasoning + 256k gives “extreme usage… performance still fantastic” — which contradicts Moonshot's own always-max guidance — and another warns Claude and Codex plans “have been heavily nerfed over the past 2 months”, so every comparison here ages fast, including ours
r/kimi · 2026-07-15 · 183 score as read 2026-07-20 (Reddit scores drift) · ⚠️ METHOD, NOT A BOAST: read via agent-reach 2026-07-20 — that render collapses reply subtrees (12
[+N more replies]nodes hiding 14 replies), so this is the top-level thread, not the full comment tree. Every quote here was matched in what rendered; none sits inside a collapsed node. 11 first-hand voices counted; 3 refused for stake (a comped beta tester, a Cline poster, and one table its own author says came from GPT)
Show all 28 sourcesShow fewer
- One evaluator, one scaffold, all three models — Terminal-Bench 2.1 on Terminus 2 / e2b, 267 instances: K3 85.0% (227/267) · Claude Fable 5 84.6% (226/267) · Opus 4.8 84.6% (226/267). K3 is one task instance ahead of each. Moonshot's “+3.7” ran K3 in its own KimiCode harness Artificial Analysis · re-fetched 2026-07-20 · Fable row is “with fallback” (Opus 4.8 fallback)
- Frontend Code Arena, blind human pairwise — Arena's own words: “Kimi-K3… is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5”, a 17-place jump from K2.6, #1 in 6 of 7 domains and “#2 only in Gaming behind Fable 5”. The only headline where K3 clears the whole Claude line — and it measures preference, not task completion. The per-model Elo table (Fable 5 1631, Opus 4.8 Thinking 1562) is not in this post — it is the relayed leaderboard, cited separately @arena · 2026-07-16 · verified verbatim 2026-07-20 from an authenticated capture · full Elo table relayed at V2EX t/1227894 #2 (@owen800q)
- Run the next day on a coding-agent repair harness against six peers: “It finished last of 7 the models with 53 of 67 attempts (79%)” — with both Fable 5 and Opus 4.8 ahead of it @AlphaSignalAI · 2026-07-17 · ⚠️ sponsor-funded, but the result is adverse to hype · verified verbatim 2026-07-20 from an authenticated capture
- You may not be able to buy it: Moonshot paused new K3 subscriptions on 2026-07-19, two days of demand having pushed its GPUs near capacity. Existing seats unaffected SCMP · corroborated by Reuters + Dataconomy
- List price, both vendors' own pages: K3 $3 in / $0.30 cached / $15 out per MTok. Claude Fable 5 $10 / $1 / $50 — exactly 3.33× K3 on every category. Opus 4.8 $5 / $0.50 / $25 = 1.67×. And K3's list price is identical to Sonnet 5's post-introductory rate Anthropic pricing docs · fetched 2026-07-20 · ⚠️ VENDOR — primary for prices only
- The access constraint most readers actually hit — and it changed on this card's own publication date, which we got wrong until a third verifier caught it. From 2026-07-20, Claude Fable 5 is a standard, permanent part of Max and Team Premium plans, at 50% of each plan's usage limits — Anthropic's own announcement, and a spokesperson on the record: “Fable 5 at 50% of usage limits is now a standard, permanent part of Max and Team Premium plans”. Max is an individual plan: $100/month (5×) to $200 (20×); Team Premium is $100/seat. Pro ($20) and Team Standard lose bundled Fable 5 — they keep access only via usage credits at API rates ($10/M input, $50/M output) after a one-time $100 credit. Two compounding catches: the 50% Fable slice is measured against limits that also fell on 07-20 when the Claude Code weekly-limit boost ended (~33% lower, the-decoder's arithmetic on the same announcement); and at $50/M output, one 2M-output-token session spends the entire $100 credit. ⚠️ THIS CARD PREVIOUSLY PRINTED “Fable 5 is ticked for no individual Claude plan at all”, in four places, reading tick-marks off a marketing table. That was FALSE on the day we published it — and it was falsified by two sources already on this card face: Simon Willison (
evidence[0]), whose next post is the announcement, and 小红書 一起 Vibe (evidence_more[21]), whose first paragraph is the announcement while we quoted her last one. The corrected claim is narrower and stronger: “Just use Fable 5” is a $100-plan answer, at half of a limit that just shrank — not a $20-plan option simonwillison.net · 2026-07-18 · relays @claudeai verbatim — “Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits” · and adds: “users on the $20/month plan will still not have access to Fable 5 on that subscription. The Max plans are $100 and $200/month” · primary: @claudeai · 2026-07-17 · ⚠️ VENDOR — the announcement itself, primary for tiers only · corroborated thenewstack · 2026-07-20 · Anthropic spokesperson on record, PCWorld · 2026-07-20 · the plan prices, the-decoder · 2026-07-18 · the ~33% baseline cut · token rate: anthropic.com/claude/fable · ⚠️ VENDOR — primary for the $10/$50 price only, and note that this page still had NOT been updated with the split when we read it 2026-07-20 - The direct subscription-to-subscription measurement, any language: “太贵了, 实测 199 的 KIMI CODE 的 K3 的量, 和 claude pro flabe 的量差不多” [sic — the source writes “flabe”] — the ¥199 K3 allowance ≈ a Claude Pro Fable allowance. 🕐 Dated: measured 2026-07-17. Claude Pro stopped including Fable 5 on 2026-07-20, so the comparator no longer exists — the reading stands as made, and cannot be re-run on a $20 Claude plan as of 2026-07-20 V2EX @m1nm13 · t/1227894 reply #40 · 2026-07-17 · 中文
- What ¥199 measurably buys, from a user's own console: ~32M tokens per 5 hours · 164M per week · 660M per month — “当然这是 199 全程使用 K3 的额度”. Moonshot publishes percentages, never token counts; Anthropic publishes neither. Both meter on a rolling 5-hour window plus weekly caps V2EX @SiWXie · t/1228031 replies #29–#30 · 2026-07-17 · 中文
- Moonshot's own concession, quoted because it's against interest — and flagged because the neutral data contradicts it: K3 “exhibits a noticeable gap in user experience compared with Claude Fable 5”. Scoped to experience, not score; the two same-harness boards show a tie Moonshot AI tech blog · ⚠️ VENDOR — a claim, and one we could not reproduce
- Moonshot's “>90% cache hit keeps cost low” defence, falsified from a user's own console: 94% hit rate, all cache hits — and two tasks exhausted the 5-hour limit. “那就是比Claude fable5还贵” (then it's dearer than Fable 5)
小红书 幸运的蜗牛 · 2026-07-17 · 75 comments · re-read 2026-07-20 via agent-reach (signed-URL read) · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring. - A Claude-side failure no leaderboard captures: on a private, un-benchmark-maxxable cybersecurity eval, “Best price/recall: Kimi K3” — best value for recall, not best recall (that is GPT 5.6 Sol, “at over 7x the price”) — while “Fable 5: 100% refusal rate. Cannot be used for security analysis.” Refusal scores zero at any price. Opus 4.8: “only recommended with subscription or high-discount API price” @cramforce (CTO, Vercel) · 2026-07-18 · verified verbatim 2026-07-20 from an authenticated capture
- The methodological warning, from a practitioner who spent $40 across three harnesses: “‘Beats Opus.’ ‘Beats Fable.’ That framing is the first thing to avoid… the same model changes rank depending on which benchmark you look at” note.com まさお · 2026-07-17 · 日本語
- The footgun Moonshot documents in its own Claude Code setup page, in a 注意 callout: “关闭 thinking 后 K3 和 K2.7 Code 都会被路由到 K2.6,请保持 thinking 开启以使用 K3 / K2.7” — turn thinking off and you are silently served K2.6, not K3, on a keystroke (⌥T / Alt+T). The same page's tier table gates context by plan: Andante and Moderato cap at
262144; only Allegretto and above get K3's1048576Kimi Code · Claude Code setup doc · re-fetched 2026-07-20 at the current URL (the oldother-coding-agents.htmlpath now 404s) · ⚠️ VENDOR - The one disposition reading in the corpus: a Chinese hands-on writeup found K3 behaves “more like a strong executor engineer than a proactively refactoring architect”, and reported 「我跑了十几个用例就烧掉了接近 30 美元」 (a dozen-odd test cases burned close to $30) getting there. Still the weakest item on this card — you cannot open the page. But it is no longer a lone capture: a third-party semantic index carries the full article text and matches both quotes verbatim, by a route with no access to anything of ours. Two witnesses to the string; still zero readable pages 51CTO 李卓 · 2026-07 ⚠ live page returns HTTP 567 (a bot-wall body) — text independently corroborated by a third-party index, 2026-07-20
- The mechanism behind the cost complaint, named for the first time — and by a route no other source in this sweep took (API, own harness, per-task console figures, no affiliate link): 「同一个简单任务(“把按钮的 border-radius 改成 8px”),K3 产生的 reasoning token 是 Claude
thinking: low模式的 6 倍」 — six times Claude's reasoning tokens to change a border-radius. His tasks billed ≈¥1.27 / ¥2.37 / ¥4.49 at 8,340 / 15,720 / 28,500 reasoning tokens. Also a specific robustness gap: K3 「没处理组件卸载的竞态条件……这个问题 Claude 基本不会犯」. His verdict line is one neither camp would write: 「K3 用 Fable 5 大约三成的价格,做到了八九成的体验」 — eight or nine tenths of the experience at about three tenths of the price 掘金 Flynt · 2026-07-19 · 中文 · found only after fixing the search syntax — this card previously published 掘金 as “returned nothing” - K3 run publicly inside Claude Code — and the uploader posts his own quota numbers. On the ¥199 plan: 「定价确实贵 199订阅演示完视频中的案例消耗了周额度的20%多」 (demoing this video's cases consumed over 20% of the weekly quota). Asked by a commenter why so little, he names the structural complaint rather than the price: 「主要是不刷新额度 要是学codex刷新额度就好了」 (the main thing is the quota doesn't refresh — it'd be good if they copied Codex and refreshed it) — the one person here with an incentive to flatter the model he is demoing, and this is his complaint. Independently, @DD蛋卷: 「实际上写后端还是很差,写前端真的强无敌,比gpt5.6sol好太多…不过kimi额度消耗是真夸张,99套餐一晚上用了1/3」. In its comments, the §harness footgun happening to a real person, unanswered: 「不懂就问,我用ccswitch把kimi接入claude code,我设定的模型是k3,为啥实际调用的是k2.6」. Top independent comment (181 likes): 「1、过度思考、效率不高…2、定价鸡肋…难以融入工作流」. ⚠️ ONE QUOTE WAS CUT FROM THIS ITEM ON 2026-07-20 — a “nearly four hours” line we could not find in the description, in 184 comments including all eight uploader replies, or in the reply API. An unconfirmable quote is treated as fabricated until proven otherwise, so it is gone, and so is the claim that rested on it B站 AI超元域 · 2026-07-17 · uploader-written replies and comments ONLY — this video has no subtitle track (the API returns an explicit empty list), so nothing spoken is quoted or characterised · method: full set is login-walled (372); pages 1–3 of the public, hot-sorted reply API were read — 107 of 372 replies, under a third, not all — and hot-sorting favours high engagement
- Someone who paid for it himself, on camera, and went back to Claude: “out of three of the four tasks I believe Fable five was in the lead and faster and included in my plan. And I didn't hit API errors… I'm still going to be pumping my money into Anthropic and Fable five” — having hit “a 403 on all of them… my usage limit literally in just a few minutes with $19”. And the rebuttal travels with it, top comment, 114 likes: “so you compared $200 Fable to $19 Kimi K3. Fair comparison, indeed.” He never states his Claude tier — which kills this as a measurement while leaving it perfectly good as a report of what $19 buys you YouTube Creator Magic · 2026-07-17 · 22,148 views · paid the $19 himself, no sponsor — unlike the two higher-reach K3 videos we disqualified (one “Sponsored by Moonshot AI”, one carrying an affiliate discount)
- The cleanest measurement of where K3's money actually goes: 「15.8万token里81%是推理,代码只占19%」 — of 158,000 tokens, 81% was reasoning and 19% was code — on a run that took 75 minutes against a rival's five. Independently corroborates the 掘金 mechanism above from a different user, a different task and a different harness. He also notes 「API定价K3是$15/M output,和Claude Sonnet 5完全同价」, which reproduces this card's own price-sheet finding. ⚠️ His cost comparator is Grok, not Claude — the 2.4× and 3×-per-line figures in that note do not transfer, and are not used here
小红书 流落人间的小丑 · 2026-07-18 · 241 likes / 106 comments · ⚠️ self-promotional framing — the numbers are first-hand, the narrative is not evidence · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring. - A Chinese practitioner arriving at this card's verdict independently, and stating it as a decision rule — which is worth more than agreement with a conclusion: 「1. 有条件用 fable5 的,不会考虑用 Kimi,平替说法不妥 2. …如果你需要编码,还需要审美,kimi 几乎是少有的良好补充 3. 如果没条件用上面 2 个…那 kimi 是个好选择」 — (1) anyone who can use Fable 5 won't be considering Kimi, and calling it a drop-in substitute is wrong; (2) if you need coding and aesthetic judgement, Kimi is one of the few good complements; (3) if you can't get either of the above, it's a good choice. That is trial-don't-migrate, second-seat and the access argument, reached in Chinese from his own use
小红书 自由之路 · 2026-07-20 · 中文 · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring. - “Open weights”, verified rather than believed: the moonshotai org's newest model is Kimi-K2.7-Code (created 2026-06-11, last modified 2026-06-15) — no K3 repo exists. Promised 07-27, licence unannounced, and zero self-hosters found in any language Hugging Face API · re-checked 2026-07-20 · machine-checkable, re-run it yourself
- The single most load-bearing item for DON'T MIGRATE, and it is an integration constraint rather than an anecdote: 「官方明确要求多轮对话和工具调用必须原样回传完整 assistant 消息,包括思考内容;中途从其他模型切换到 K3 也可能导致生成质量高度不稳定,因此并非所有现有 Agent 框架都能无缝替换」 — the vendor explicitly requires that multi-turn conversation and tool calls echo back the complete assistant message verbatim, including the thinking content; switching to K3 mid-stream from another model can make generation quality highly unstable; so not every existing agent framework can swap it in seamlessly. This is the §harness-coupling risk stated plainly by a third party — the thing both our switch-if and stay-put columns turn on
小红书 老K · 2026-07-20 · 9 likes · 中文 · durable key 「并非所有现有 Agent 框架都能无缝替换」 · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring. - A SUSTAINED-USE report — and it is the strongest first-hand statement on this card that Fable 5 is genuinely ahead. We are giving it real placement precisely because it cuts against the model we are otherwise defending from bad benchmarking: 「但有一说一,经过最近的高强度使用(不想浪费重置),Fable 5 还是显著优于 Kimi k3 和 GPT 5.6 Sol Ultra 的,无论是聪明、灵巧、agentic coding,就是有一股灵气、毫不费力超出预期完成任务的感觉,其他模型还不能替代…」 — after recent high-intensity use, Fable 5 is still markedly better than Kimi K3 and GPT-5.6 Sol Ultra, on intelligence, dexterity and agentic coding; there is a certain spark to it, and the other models still can't replace it. ⚠️ This card previously printed “no sustained-use report exists”. That was false — three surfaced across a 33-note sweep, and this is the one that hurts
小红书 一起 Vibe · 2026-07-18 · 11 likes · 中文 · durable key 「就是有一股灵气」 · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring. - The counter-example that killed one of this card's own claims. We wrote, on this card face, that in the Chinese corpus nobody is arriving from Claude. One person is: 「说实话不是一直好用,几个月前的kimi code还是很渣的。但最近突然好使了起来,好使到我直接把claude退订了」 and 「个人感觉,最近kimi cli的表现已经超过claude了」 — “good enough that I cancelled my Claude subscription outright” / “Kimi CLI's recent performance has already surpassed Claude's”. So the absolute is withdrawn and replaced by a count, which is the stronger claim anyway: across 33 小红书 notes, exactly one first-hand Claude→Kimi cancellation surfaced, while the dominant exit vector from Claude is Claude → Codex, with Kimi a secondary seat. ⚠️ Carried with its own caveat: 2 likes, and the note's last line has a genuinely ambiguous referent we refuse to resolve by guessing — it does not touch the self-contained Claude sentence
小红书 Jake · 2026-07-20 · 2 likes · 中文 · durable key 「好使到我直接把claude退订了」 · signed link, verified live 2026-07-20 — tokens expire: if it 404s, search the author or the durable key and open the signed link the search returns; a bare
/explore/<id>is rejected. Durable key = note id + author + body substring.
Why confidence is LOW
Coverage — what we read, what we skipped
⚠️ THIS FOOTER WAS REWRITTEN ON 2026-07-20 BECAUSE MOST OF IT WAS WRONG — and wrong in our own favour's opposite direction. Rounds 1–3 of this card were swept with plain web fetching. The research tool our own sweep contract mandates (agent-reach) was never called. The result was not a card that over-claimed. It was a card that invented weaknesses it does not have: four 小红书 sources were published as “login-walled, not re-verifiable” when they read perfectly well (小红书 needs a signed URL; we kept fetching the unsigned one, got a JavaScript shell, and called it a login wall); B站 was characterised without ever being read; 掘金 was written off on a search-syntax artefact; and 知乎 was described as empty when in fact its search is CAPTCHA-walled. A card that invents a weakness is as dishonest as one that invents a strength — it just feels like modesty at the time. Everything below is what the re-sweep actually found.
READ: Eval houses — Artificial Analysis (Terminal-Bench 2.1 same-harness, model comparison page, GDPval-AA v2), DeepSWE/Datacurve, tbench.ai (as a guard, not a number) · GitHub primary bug reports via the API (hermes-agent #67386, opencrabs #616, opencode #37635 + a five-issue prior-generation cluster, kimi-cli #1994) · Hugging Face API + K2.7 licence text · Vendor pricing pages, both sides (platform.claude.com pricing docs, claude.com/pricing, Anthropic Pro/Max support articles, Moonshot's blog + three doc sites) — ⚠️ read as primary for prices and tiers only, scored neutrality 0.0–0.2 in the ledger, never allowed to carry a verdict · Reddit (r/ClaudeCode, r/kimi ×2 incl. the 60-comment subscription-comparison thread, r/codex, r/OpenAI, r/hermesagent) · Hacker News — three threads, cited separately with real counts: 48935342 1,206 comments / 2,090 pts, 48960218 598 / 623, 48947717 221 / 403 · X practitioners · simonwillison.net ×2 posts (the K3 post AND the 07-18 Fable-permanent post — the second added only in round 5, see below) · Trade press on the 2026-07-20 Anthropic plan change (thenewstack, PCWorld, the-decoder, techtimes, androidauthority, aipricing.guru) plus @claudeai's own announcement · 中文: V2EX ×4 threads, 小红书 ×33 notes (this READ list said ×9 until 2026-07-20 — the pre-round-5 figure, never updated when the saturation sweep tripled it, while three sentences further down this same footer called it a 33-note sweep. 9 is 33 minus the 24 that were new. Corrected under this card's own published rule: an inaccurate method statement is inaccurate in either direction.), 阅微堂 (zhiqiang.org), 人人都是产品经理, 知乎, 掘金, B站 ×5 uploaders read of 13 found · 日本語: note.com, Qiita, PC Watch · YouTube + transcripts · both Google SERPs.
WHAT THE LEDGER BELOW CONTAINS, AND WHAT IT DELIBERATELY DOES NOT. Every one of the 65 voices behind n=65 now has a ledger row, which is the point of the ledger and was not true until 2026-07-20: 14 of the 65 were uncountable from this page — 13 小红书 notes from the saturation sweep with no row at all, plus one roster line citing a price instead of its note id — and the roster that lets you recount n is in the receipts, which this site does not publish. So the ledger was the only audit surface a reader had, and for 14 voices it was empty. Each of the 13 was re-found through agent-reach (search → signed URL), re-read end to end, and ledgered with its note id, author, a distinctive body substring and its weight; a further 5 roster lines pointed at ledger ids that did not exist and now cite the real ones. Five items this sweep read are deliberately NOT in the ledger, named here so the omission is auditable rather than silent: two 小红书 notes that are image-only — the body is a picture, so nothing is quotable from them (「Kimi K3 风评大变」 6a5c99bd…, 102 likes; 矩元 6a59ee33…, 9 likes) — and three notes dated 2026-07-21 (6a5e4cea…, 6a5e5318…, momo 6a5e50c6…), quarantined because a card dated 07-20 may not cite 07-21, which is a rule about principle and not about their merit. They are listed so the next sweep collects them instead of rediscovering them.
CORRECTED ON THE RE-SWEEP (2026-07-20) — every one of these was a claim this card made, and got wrong: (1) The four 小红书 “not re-verifiable” flags are GONE. All four notes were read via agent-reach 2026-07-20 (search → signed URL), and each was read end to end — these are single-page notes with no pagination and no collapsed replies, which is why an exhaustiveness claim is safe here and was not safe on the Reddit thread. To re-check one yourself: search the author or a phrase, then open the signed link (?xsec_token=…) the search returns — a bare /explore/<id> is rejected. Signed links expire, so for each note the receipts record the durable key instead: note id + author + a distinctive body substring. (2) A quote of ours was not verbatim. 终梦's conclusion was printed inside 「」 under a heading reading Verbatim as 「日常杂活 DeepSeek,最难的活给 K3」. He wrote 「我的方案:日常杂活 DeepSeek 焊死,最难的活给 K3 单独开小灶。」 A tidied-up abbreviation presented as a quotation is a fabricated quote that happens to be true; it is fixed everywhere and logged with the same weight as this card's round-1 fabrication. (3) “Three live session-wedging bugs in three days” was one bug. opencrabs #616 is closed — by the repo owner who filed it, nine minutes later; opencode #37635 is an opencode-go gateway bug hitting fifteen models with kimi-k3 as one list entry. Only hermes-agent #67386 is live — and on re-verification it is not K3-specific either: its body opens “Strict OpenAI-compatible providers (Moonshot/Kimi, ZAI)…” and it carries provider/kimi and provider/zai, so it is a two-provider wire-format class bug, not a K3 defect. Zero of the three original bugs is both live and K3-specific. On its labels: 6 of the 9 are ordinary repo taxonomy (type/bug, comp/agent, provider/kimi, provider/zai, P2, area/sessions) and only 3 carry the sweeper: prefix, so “maintainer-labelled” is now “repo-labelled”. kimi-cli #1994 was dated 2026-06; it was filed 2026-04-22. (4) There is no “~1,850-comment HN thread”. That number was two separate threads added together (1,206 + 598) and then described as one. Three real threads, listed above. (5) “Every V2EX reply re-read line-by-line” was not true. V2EX's API has an SSL failure, so those threads come through Jina Reader — whose render is lossy and silently drops replies (eleven missing from t/1228031 alone). Every V2EX quotation on this card did check out; the claim of exhaustive reading did not. (6) Two comment counts were stale (幸运的蜗牛 59→75, blueraincoat 36→49) and one author was misattributed: “V2EX @blueraincoat” — there is no V2EX evidence for him at all; he is 小红书 only. (7) 阅微堂 was fine all along — it is simply the blog name for zhiqiang.org, which has been in the ledger with a live URL throughout. It read as unfindable only because this footer printed the Chinese name with no domain beside it. Named properly now, not deleted.
UPGRADED, NOT DOWNGRADED: x.com is no longer login-walled to us. All four x.com items were re-verified verbatim through agent-reach, independently of the operator's browser capture — so they are now re-verifiable by you, not merely “captured from a session we hold”. The re-check also confirmed the card's own caution: the @arena post contains 1679 and “surpassing Claude Fable 5” and contains neither 1631 nor 1562, exactly as the card says. · 51CTO now has two witnesses. Its page still returns HTTP 567 (a bot-wall body; agent-reach gets an empty body) — but a third-party semantic index carries the full article text and matches both quoted strings verbatim, by a route with no access to anything of ours. It is the only genuinely hard-to-check source left on this card, and still its weakest item.
GAPS, STATED HONESTLY AND QUOTED WHERE WE HAVE THE ERROR: 知乎 enumeration was never achieved. Article pages read fine; zhihu.com/search returns 「系统监测到您的网络环境存在异常」 and “This page maybe requiring CAPTCHA”. So we can say what the pages we reached contain — the one substantive item was Moonshot's own launch blog, reposted, which carries nothing — but not what 知乎 contains. That is narrower and truer than the “returned nothing” this card used to print. · 微信公众号 via Sogou remains inaccessible — a real gap, disclosed. · B站 spoken content is not cited anywhere on this card. Both videos have no subtitle track — the player API answers code 0 (OK) with an explicit empty subtitle list, which is a definitive negative, not a fetch failure. So every B站 item here is existence, uploader-written description, or comment, never speech. And the largest video's full comment set is login-walled (「登录后查看 372 条评论」): pages 1–3 of the public, hot-sorted reply API were read — 107 of 372 replies, under a third, not all of it, and hot-sorting favours high engagement. (This footer previously said “the twenty comments read”, which under-reported the sweep and implied a ceiling never tested. An inaccurate method statement is inaccurate in either direction.) · Zero self-hosters of K3 in any language. · Zenn, はてなブックマーク and Korean (Velog/tistory) had no hands-on K3 coding writeups. · The “excessive proactiveness” check still finds nothing — five phrasings through an authenticated Reddit session returned no on-topic result, and the three HN threads grepped for proactive/improvis/overeager return zero. (Stated that way deliberately: with semantic search a non-match comes back full of unrelated posts, so the old “empty result set” overstated how clean the null was.) Meanwhile it has two first-hand contradictions that survive a stake check, and zero corroborations — one of them a user documenting K3 retracting its own conclusion unprompted. ⚠️ It was three until 2026-07-20, and we removed one ourselves: the ShipSolo production writeup is real and readable, but its author's own second paragraph discloses pre-release access held under embargo (「发布之前,我已经把 ShipSolo 微信小程序这个真实项目整个交给它」). That is the same stake class as the comped beta tester we already exclude by name, and it was being used to rebut a vendor's own disclosed weakness in that vendor's favour — the highest-risk seat an unaudited source can occupy here. It is kept in the ledger with its URL and its stake, and carries nothing on this card. Neither surviving contradiction compares against Claude: Z3PH1NUE's note writes only 「5.5」 and does not disambiguate it — on context that reads as GPT-5.5, which we label an inference; what is certain is that it is not a Claude model.
REFUSED, AND WHY (iron rule #1 — readers pay, vendors never do): AICodeKing (YouTube) discloses “Sponsored by Moonshot AI” — one of the highest-reach English K3 “tests”, and it carries nothing here. Dubibubi carries an affiliate offer for the product under review. 小红书 Kimi智能助手 is Moonshot's own account, and its launch post (8,022 likes read 2026-07-20; it was 8,016 when first captured, and the figure is a timestamp rather than a constant) is the highest-engagement K3 item in the whole 小红书 corpus — which is the concrete reason engagement volume is used as a signal nowhere on this card. Three Reddit commenters were excluded from the count for stake: a comped Kimi beta tester, someone who posts for Cline, and one whose comparison table his own comment says came from asking GPT. The page-one SEO/reseller cluster (kie.ai, myclaw.ai, layer3labs.io, llm-stats.com, codersera, vallettasoftware, digitalapplied) was read for cross-checkable facts and discounted for the verdict: 4 of 10 page-one results for “Kimi K3 vs Claude” and 6 of 10 for “vs Claude Fable 5” are commercial or affiliate properties, and they uniformly republish Moonshot's KimiCode-harness 88.3 as a like-for-like win — the most propagated error about this model. Named 小红书 users allege the pro-K3 wave was bought: 「Kimi K3凌晨上线,随后就是铺天盖地的商单,都快把真实用户的声音淹没完了」, seconded independently by another. We cannot verify that; we can and do weight bug reports and same-harness numbers above sentiment volume because of it.
WATCH OUT: The number this card most wants still does not exist — but the silence is smaller than we said. No independent test has run the same K3 tasks in both Kimi Code and a third-party harness, so the harness penalty itself is still unquantified — though K3 has now been run publicly inside Claude Code (AI超元域, B站, 2026-07-17) and inside a shared agent harness alongside rivals (Token就是词元, B站, 2026-07-18). Neither closes it: the Claude Code run used different tasks on the Kimi side, the shared-harness run isolates the model, not the harness, and the one YouTube head-to-head moved two variables at once (model and harness, plus an undisclosed Claude tier against a $19 plan). · The ¥699 Kimi-client price is contested — one first-hand user says ¥699 for 1M context, another the same week says ¥199, and nothing adjudicates. Only the Kimi Code tiering is corroborated by both (¥199 = 1M · ¥99 = 256k · ¥49 = no K3). · The “max” in the benchmark rows is a reasoning-effort setting, not the Claude Max plan — and the Fable row on Terminal-Bench is additionally an Opus 4.8-fallback configuration. Both same-harness gaps against Fable 5 are smaller than their own error bars and point in opposite directions; treat “K3 beats Fable” and “Fable beats K3” as equally unsupported. · Sustained maintenance on a large existing repo is unmeasured in every language. · And read the direction before you read the verdict — as a COUNT, because the absolute we used to print here was FALSE. This footer said “nobody is arriving from Claude”. One person is (小红书 Jake, 「好使到我直接把claude退订了」, cited above), and a second arrives after an account ban with Codex as his destination. What a 33-note sweep actually supports: the dominant exit vector from Claude in the Chinese corpus is Claude → Codex, with Kimi a secondary or supplementary seat, and exactly one first-hand Claude→Kimi cancellation surfaced across those 33 notes. Most arrivals come from Codex, GLM and DeepSeek; one recommends 「100刀gpt pro, 20刀claude,gpt当主力」; one says Claude and GPT resetting allowances makes him 「真的更想直接去用它们」. Claude is the seat these practitioners are trying to keep, not the one they are leaving — which is what the ⇄ is for.
⚠️ WHAT ROUND 5 REMOVED FROM THIS CARD, after TWO independent adversarial verifiers both refused to pass it. Both agreed the verdict, the arithmetic and the vendor hygiene were sound. Both refused the same thing: one recurring defect class — asserting a characterisation of material that was not read, or cannot be reached. That is bug #39's exact shape, and it was committed several more times inside the round convened to close it. (1) “Read in full” on the Reddit thread was FALSE — that render collapses 12 reply subtrees; the claim is replaced everywhere by a method and its limits. (2) A B站 quote could not be located in the description, in 184 comments including all eight uploader replies, or in the reply API — it is deleted and treated as fabricated, along with the claim resting on it. (3) Two universal negatives on this face were falsified by a 33-note saturation sweep — “nobody is arriving from Claude” and “no sustained-use report exists”. Both are now counts. (4) Our own cuts register was certifying a TRUE sentence as fabricated (「官方建议64卡起步,普通玩家基本不用想」 is 李卓's, on 51CTO, a source we already cite). A false accusation of fabrication, living inside the artifact whose whole job is to prevent fabrication, is worse than the original error — which was a misattribution. (5) Three 小红书 links on this face did not open, in the exact form this footer tells you is rejected. Every 小红书 item now carries a signed link plus the search-then-sign instruction. (6) n=75 could not be audited — there was no per-voice roster anywhere, so nobody could recount it. There is now a 65-row roster with a stated rule, and the number went DOWN, 75 → 66 → 65. No evidence was deleted; nine analysts and relays simply stopped being counted as voices. And one number in this round's own brief was wrong: we were told the Reddit thread stood at 187 and one commenter at 20; the live source read 183 and 19. We followed the source.
⚠️ ROUND 5 WAS THEN FAILED BY A THIRD INDEPENDENT VERIFIER, AND THIS IS WHAT HE FOUND — the worst single defect in this card's history. Four claims on this card face were FALSE ON THE DAY WE PUBLISHED THEM. We wrote, in the verdict, in a switch-if line, in a dropdown item and in the compatibility grid, that Fable 5 was included on no individual Claude plan at all. On 2026-07-20 — this card's own date — Anthropic made Fable 5 a permanent part of Max and Team Premium at 50% of usage limits, and Max is an individual plan starting at $100/month. Pro ($20) and Team Standard lost bundled access the same day and moved to usage credits at $10/M in, $50/M out after a one-time $100 credit; and the 50% slice is measured against a weekly baseline that also fell that day (~33%, the-decoder). It is corrected in all four places, and the corrected claim is sharper than the overreach: Fable 5 needs a $100+ Max plan, at half of a limit that just shrank. Here is the part that should cost us your trust rather than earn it back: the announcement was carried by two sources already on this card face. Simon Willison is evidence[0] — we quote his 07-16 K3 post and never opened his 07-18 post, which is the announcement. And 小红書 一起 Vibe is evidence_more[21], whose note is literally titled 「Fable 5 永不下线」 (“Fable 5 never goes offline”) — we quoted its last paragraph and never read its first, which opens 「Claude 果不其然怂了,宣布从7月20日起,Fable 5 永久在 Max 和 Team Premium 计划中,但限额仍然是50%」. Grepping all four receipt files for Team Premium and 永久在 Max returned zero hits. This was never weighed and rejected. It was never seen. A quote extracted from a page is not the same thing as a page that was read, and that is now the standing lesson of this card. The knock-on, stated plainly rather than buried: this card's strongest quantified finding — u/djdante's $20-vs-$20, 4× weekly quota — was measured 2026-07-15 against a Claude Pro tier that included Fable 5 and no longer does. It is not deleted: it was valid when made and a verifier re-read it verbatim at source. It is date-stamped everywhere it appears, as is u/ahnyudingslover's $20-tier corroboration and @m1nm13's “¥199 ≈ a Claude Pro Fable allowance”. And one more over-claim came out of the roster in the same pass: 小红書 老K was counted as a first-hand voice; his note is a 汇总, a compilation of Artificial Analysis data with no hands-on use, which the roster's own rule excludes. n: 66 → 65. Read our confidence with all of this in view: it stays LOW, and the reason is now sharper than it was — the defect landed on the subscription-economics half, the half this card calls “close to settled”. Near-settled describes a moving object. Check the plan tables yourself before you spend anything.
Show every source we read (111) Hide the ledger
The 111 sources in this card's ledger — 28 of them quoted above. Sources we read and did not quote are listed too, with why. This is the ledger, not a claim of exhaustiveness: anything the sweep read but deliberately left out of it is named in the coverage note above, with the reason. Reliability and neutrality are our own scores, not the publisher's. Vendor-owned pages are marked in the notes and never carry the verdict.
- personal blog — simonwillison.net/2026/Jul/16/kimi-k3/ THE crux source for this card. States plainly that K3 mostly beats Opus 4.8 while LOSING to Fable 5 — i.e. the answer flips depending on which Claude you're on. Also the primary for the 25c pelican, the always-max reasoning cost, the 85-token hidden system prompt, and the Anthropic tokenizer asymmetry. No vendor stake; discloses his sponsors.
- vendor blog — www.kimi.com/blog/kimi-k3 VENDOR — every benchmark claim here is a CLAIM, not a finding. Quoted only for admissions against interest: the thinking-history instability warning, excessive proactiveness, and the concession that K3 trails Claude Fable 5 and GPT-5.6 Sol. Its Terminal-Bench 88.3 was run in Moonshot's own KimiCode harness and is NOT comparable to the 84.6 it cites for Opus from a third party.
- eval house — artificialanalysis.ai/models/comparisons/kimi-k3-vs-claude-opus-4-8 Fetched directly during the sweep. Intelligence Index K3 57 vs Opus 4.8 56; blended price $2.31 vs $3.85; TTFT 1.99s vs 40.49s. Notably lists BOTH as 'Open Source (Weights): No' — AA classifies K3 as proprietary because the weights are not out yet. Independent of both vendors; sells data services, which is a mild commercial stake.
- eval house — artificialanalysis.ai/evaluations/terminalbench-v2-1 One of only two same-harness numbers that exist. BOTH models on Terminus 2, e2b sandbox, 267 instances: K3 85.0% vs Opus 4.8 84.6% — one task instance apart. This is what collapses Moonshot's +3.7 claim into a tie. Measures Opus, NOT Fable.
- eval house — deepswe.datacurve.ai/ The other same-harness number, and the one real reproduced K3 win over Opus 4.8: 69% ±5% vs 59% ±2%, CIs clear of each other, $4.65 vs $13.22 per task, on FEWER output tokens. Board note: 'All models run on mini-swe-agent for consistency.' Moonshot UNDER-claimed this one. Measures Opus, NOT Fable.
- eval house — www.tbench.ai/leaderboard/terminal-bench/2.1 Read as a guard, not for a number. The official PR-gated board does NOT list K3 at all and puts Opus 4.8 + Claude Code at 78.9% ±1.3% — ~6 points below AA's figure for the same model on the same benchmark. Different scaffolds. Quoting one model from one board against another model from the other board manufactures a gap out of nothing.
- eval house — artificialanalysis.ai/evaluations/gdpval-aa GDPval v2 — the largest Claude-side lead on any benchmark. ✅ FIGURES SETTLED 2026-07-20 against the PRIMARY source (Artificial Analysis' own launch post, authenticated capture): Claude Fable 5 1760 · Kimi K3 1668 · Claude Opus 4.8 1600 · GLM-5.2 1514 · GPT-5.5 1494 · Kimi K2.6 1190. K3 sits BETWEEN the two Claudes (+68 over Opus 4.8, -92 to Fable 5) — the same shape as every other axis on this card. The earlier hedge ('sources disagree: 1668/1685/1687 and 1760/1815') is withdrawn; the 1687 / 1815 / 1683.65 variants were third-hand relays and are dropped everywhere.
- x — x.com/ArtificialAnlys/status/2077832874183860404 PRIMARY for the launch measurements, and now the AUTHORITATIVE source for the GDPval v2 Elos: K3 1668, Claude Fable 5 1760, Claude Opus 4.8 1600, GLM-5.2 1514, GPT-5.5 1494, K2.6 1190 — which settles a figure the card previously refused to state because relays disagreed. Also: AA Intelligence Index 57 ("comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol"); cost per task $0.94 vs Opus 4.8 $1.80; 21% fewer output tokens than K2.6; AA-Briefcase Elo 1547, behind only Fable 5; AutomationBench-AA 53% (#1); first-party API $3.00/$15.00 per 1M in/out with cached input at $0.30. Note AA publishes MULTIPLE cost indices that disagree, so any AA cost figure must name its index. ✅ VERIFIED 2026-07-20 from the authenticated capture (captures-x-2026-07-20.md).
- x — x.com/arena/status/2077824029126504525 Frontend Code Arena: K3 #1. ⚠️ SCOPE CORRECTED 2026-07-20 against the authenticated capture — the post itself says "Kimi-K3 ... is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5", a 17-place jump from K2.6, #1 in 6 of 7 domains and "#2 only in Gaming behind Fable 5". It does NOT contain the per-model Elo table: Fable 5 1631 and Opus 4.8 (Thinking) 1562 come from the RELAYED leaderboard at V2EX t/1227894 reply #2 (@owen800q), and the card now cites them there rather than to this post. Trusted as a real blind-human preference Elo — but it measures which output people PREFER LOOKING AT, not task completion, and the card says so. ✅ VERIFIED 2026-07-20 from the authenticated capture (captures-x-2026-07-20.md).
- trade press — www.scmp.com/tech/article/3361172/kimi-k3-developer-suspends-new-subscriptions-amid-compute-constraints Moonshot PAUSED NEW K3 SUBSCRIPTIONS on 2026-07-19 — GPUs near capacity after two days. Decision-ending for many readers: you cannot migrate onto a seat you cannot buy. Corroborated by two further independent outlets below, which is why it carries despite being press rather than primary.
- tech press — dataconomy.com/2026/07/20/moonshot-ai-kimi-k3-subscribers-halted-servers/ Corroboration #2 for the subscription pause. Read, not quoted — SCMP is the better-sourced version of the same event. Listed because corroboration is why the claim is on the card at all.
- wire (Reuters syndicated) — finance.yahoo.com/technology/ai/articles/chinas-moonshot-pauses-kimi-subscriptions-080317250.html Corroboration #3 for the pause, and adds the membership restructure (Kimi Web/App/Work split from a separate Kimi Code Membership). Read, not quoted directly.
- github — github.com/NousResearch/hermes-agent/pull/67386 THE Claude-specific footgun, and the ONE genuinely live session-wedging bug that names Kimi at all — and NOT a K3-specific one: sessions permanently wedged, hit four times in one day on kimi-k3 via OpenRouter, with one entry path being specifically Anthropic-style content thinking-blocks stripped by the wire converter. Code, logs and a fix — not an opinion. ⚠️ RE-CHECKED VIA THE GITHUB API 2026-07-20 AND DOWNGRADED IN DESCRIPTION: state is genuinely OPEN and third-party — BUT it is NOT K3-SPECIFIC: the body opens 'Strict OpenAI-compatible providers (Moonshot/Kimi, ZAI)...', the PR is titled 'fix: Moonshot/Kimi assistant must not be empty session wedge', and it carries provider/kimi AND provider/zai. It is a two-provider wire-format class bug that K3 users hit, not a defect unique to K3 — the same discount this card applies to opencode #37635. Anthropic-style thinking-blocks are ONE of four enumerated entry paths, not the cause. On labels: 6 of the 9 are ordinary repo taxonomy (type/bug, comp/agent, provider/kimi, provider/zai, P2, area/sessions) and only 3 carry the `sweeper:` prefix, so an earlier draft's 'bot-applied' overstated it. The author is an OUTSIDE CONTRIBUTOR ON A FORK, and the body is marked '🤖 Generated with Claude Code'. The card said 'maintainer-labelled provider/kimi, P2'; it now says 'repo-labelled'. A sweeper-prefixed label is not a maintainer's judgement, and the difference matters because this item is doing real work on the card.
- github — github.com/adolfousier/opencrabs/issues/616 Log-backed: Moonshot's CODING endpoint emits zero reasoning_content deltas across a full day and leaks reasoning as ordinary content, while its API endpoint separates it correctly. ⚠️ HEAVILY DOWNGRADED 2026-07-20 AFTER A GITHUB API CHECK — THIS IS NOT A LIVE BUG. `state: closed`, `state_reason: completed`, `author_association: OWNER`. Opened 2026-07-19T12:56:06Z and closed BY ITS OWN FILER at 2026-07-19T13:05:32Z — NINE MINUTES. A repo owner filing and immediately resolving an issue in his own harness is a development note, not a report of a defect anyone else hit. The card counted it toward 'three live session-wedging bugs in three days'; that claim is withdrawn (see round4_note). The underlying observation about Moonshot's two endpoints disagreeing may well be true, but this source cannot carry it as a live third-party bug.
- github — github.com/anomalyco/opencode/issues/37635 Gateway-level reasoning_content/content inversion; agent loop exits after step 1; kimi-k3 is in the affected list. Weighted DOWN because it hits all opencode-go models, not K3 specifically — included so the card doesn't overstate K3-specific breakage. ⚠️ RE-CHECKED VIA THE GITHUB API 2026-07-20: state OPEN, but the scope is even wider than the card allowed for — it is an opencode-go GATEWAY bug hitting FIFTEEN models, and kimi-k3 is ONE LIST ENTRY among them. It is not K3 evidence; it is evidence that a gateway is broken for everyone. The card no longer counts it as one of the K3 session-wedging bugs.
- github — github.com/anomalyco/opencode/issues/10996 Same failure class across three Kimi generations and four harnesses: opencode #10996 / #29690 / #25001, zed #51743, llama.cpp #20008. Read as precedent — establishes the thinking-history problem predates K3, which matters because K3 was explicitly TRAINED to depend on that echo-back. Not quoted individually.
- github — github.com/MoonshotAI/kimi-cli/issues/1994 Filed on Moonshot's OWN repo, so it cannot be curated away: quota opacity and severe throttling, K2.6 era. High trust, LOW recency — read for pattern context only, NOT cited as K3 evidence, which remains the correct handling. ⚠️ DATE CORRECTED 2026-07-20 via the GitHub API: `created_at` is 2026-04-22. This entry recorded '2026-06', which is the UPDATE date — a three-month drift on a source whose whole value is its date. Being K2.6-era is the entire reason it is not cited as K3 evidence, so getting its age wrong mattered even though the handling was right.
- huggingface — huggingface.co/api/models?author=moonshotai The cleanest fact on the card: newest model in the moonshotai org is Kimi-K2.7-Code (2026-06-15). NO K3 repo exists. Anyone can re-run this call. This is how 'open weights' is verified rather than believed.
- huggingface — huggingface.co/moonshotai/Kimi-K2.7-Code/raw/main/LICENSE PRECEDENT ONLY — this is K2.7's licence, not K3's, which is unannounced. Modified MIT with a >100M MAU / >$20M monthly revenue UI-attribution clause. Not OSI-clean. Listed so 'open source' is not taken at face value.
- vendor docs — platform.claude.com/docs/en/about-claude/pricing ⚠️ VENDOR (Anthropic) — PRIMARY AND AUTHORITATIVE FOR PRICES, and permitted to carry nothing else. Re-fetched 2026-07-20. Fable 5 $10 in / $1 cache read / $50 out per MTok; Opus 4.8 $5/$0.50/$25; Sonnet 5 $3/$0.30/$15 from 2026-09-01 — identical to Kimi K3's list price. Yields the exact ratios the card cites (Fable = 3.33x K3 on every category, Opus = 1.67x). Also the source for the tokenizer disclosure, quoted because it is AGAINST INTEREST: Opus 4.7+, Fable 5, Mythos 5 and Sonnet 5 'produce approximately 30% more tokens for the same text', which means a per-MTok table UNDERSTATES Claude's real cost per unit of work. neutrality 0.1: a vendor quoting its own price list is primary; a vendor quoting its own quality is not.
- x — x.com/kunchenguid/status/2078196563084824627 Highest-value independent voice on subscription economics, and one of very few that names Fable directly: 'Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan.' Builds a competing coding agent (myfirstmate) but has no stake in either K3 or Anthropic. Also the strongest instruction-following complaint. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- personal blog — stephen.bochinski.dev/blog/2026/07/18/the-kimi-k3-moment/ The strongest pro-K3 first-hand report ('I can't tell them apart'), and deliberately paired against @kunchenguid who ran the same comparison and found the opposite. Neutrality docked: the post is heavily political about US AI policy and ends 'I can't come up with a reason to keep paying for Claude' — and it never names WHICH Claude. Quoted for the observation, not the conclusion.
- x — x.com/AlphaSignalAI/status/2078172629496746183 Sponsor-funded newsletter, so neutrality is docked — BUT the result is adverse to hype (K3 last of 7 on a repair harness, behind BOTH Fable 5 and Opus 4.8), which is exactly the direction a sponsored outlet has no incentive to publish. Weighted up on that basis. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- x — x.com/cline/status/2078571637348372625 VENDOR (Cline is a harness shipping K3 support) — framing discounted, numbers retained. Nearly the only same-harness A/B against FABLE rather than Opus: K3 1.7x more tokens, Fable 3.4x faster, K3 2.3x cheaper. n=1 task. Included because the Fable axis is otherwise almost empty. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- x — x.com/cramforce/status/2078574147333152957 Vercel's CTO on a private, un-benchmark-maxxable cybersecurity eval: K3 best price/recall, above Opus 4.8, with Fable 5 at a 100% REFUSAL rate. Neutrality docked (Vercel has AI-product interests) but the eval being private makes it un-gameable, and refusal is a real Claude-side failure mode no leaderboard captures. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- x — x.com/QianyiZhan95891/status/2078483600815927629 avante.nvim's author asking the field the correct question — official Kimi Code, or your own agent? — plus a reply noting Claude Code once published a degradation report about over-aggressive thinking truncation. Reliability capped: recollection, not measurement. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- x — x.com/deredleritt3r/status/2078606700496773166 'K3 is also significantly more inconsistent than Gemini 3.1 Pro… struggled with fairly easy questions.' Read; contributes to the inconsistency cluster rather than being quoted alone. Comparator is Gemini, not Claude. ✅ VERIFIED 2026-07-20: captured verbatim from the operator's AUTHENTICATED browser session (x.com is login-walled to servers) and pinned in cards/kimi-k3-vs-claude/captures-x-2026-07-20.md. Every quotation on the card was re-checked against that capture. This source is NO LONGER an unverifiable item.
- reddit — www.reddit.com/r/ClaudeCode/comments/1v0cj5q/ The strongest frontend anecdote, and notable for WHERE it is — a 9-year fullstack dev praising K3 inside Claude Code's own subreddit, which is a hostile venue for the claim. Reliability capped: single greenfield POC, self-reported, no artefact.
- reddit — www.reddit.com/r/kimi/comments/1v0d2hd/ English-side corroboration of the Chinese quota finding on a different tier: $19 plan, ~40 min of work = 100% of the 5-hour limit = 20% of weekly, and that tier caps context at 256k not 1M. Matters because it shows the quota complaint is not a China-only or a currency artefact.
- reddit — www.reddit.com/r/codex/comments/1uyj6pq/ The best POSITIVE long-session report found: 1.5 hours, 8 subagents, full project review in 180k of context, 'without a single agentic mistake'. Read and counted; deliberately set against the negative long-session reports rather than cited as proof.
- reddit — www.reddit.com/r/OpenAI/comments/1uyo1hr/ Source of the sharpest openness line in the sweep — 'not open harness, ie just because you get the weights doesn't mean you'll get the same quality locally' — which is the link between the weights story and the harness story.
- reddit — www.reddit.com/r/hermesagent/comments/1uyr7k7/ Read as a TRUST signal, not evidence about the model: multiple independent users across r/hermesagent, r/OpenAI and r/codex alleging coordinated posting. Directly justifies weighting GitHub bug reports and eval boards above sentiment volume. Same thread also carries a 'context rot after 300k' report. ✅ ROUND 4: the allegation now has NAMED FIRST-PARTY sources rather than only English pseudonyms — see xhs-dabuka.
- reddit — www.reddit.com/r/kimi/comments/1uz6bll/ ★ THE MOST IMPORTANT FIND OF THE ROUND-4 RE-SWEEP. 'How does Kimi Code subscription compare to Claude Code / Codex?' — 183 score as read 2026-07-20 (Reddit scores are live and drift). METHOD: read via agent-reach 2026-07-20; the render COLLAPSES REPLY SUBTREES (12 '[+N more replies]' nodes hiding 14 replies on that read), so this is the top-level thread plus rendered first-level replies, NOT the full comment tree. WHY IT MATTERS: the card's most load-bearing claim — that K3's API cost advantage does NOT survive onto a subscription — rested on THREE voices, one of which the card had wrongly published as unverifiable. This one thread takes it to ~15 re-verifiable first-hand reports and supplies the first QUANTIFIED measurements anyone has published. THE TWO QUANTIFIED: u/djdante (L0, score 7) — ⚠️ STAKE DISCLOSED IN ROUND 5, in the direction that COSTS this card: his own first sentence is 'I ran a bunch of tests yesterday FOR MY YOUTUBE CHANNEL'. No artefact, no harness, no task set is published anywhere. He is NOT disqualified — an interest in publishing is not a vendor payment and he is comped by neither side — but the single most load-bearing number on this card is one person's unaudited recollection, and the card now says so. — '$20 Kimi plan and a $20 Claude plan - Kimi k3 will consume 4X more of your weekly quota than Fable 5 High for the same task'; u/ahnyudingslover (⚠️ L1 — a DIRECT REPLY to djdante, score 2, opening 'I can attest to this as well'; this is CORROBORATION IN-THREAD, NOT an independent measurement, and the word 'independently' is withdrawn wherever this ledger used it of him) — '3 different exact prompt comparisons between fable5 and Kimi k3 on a production codebase - fable5 uses about 30-50% of my 5 hour usage limit on the 20usd plan for both. Output quality was close, which was shocking.' READ TOGETHER THOSE TWO ARE THE WHOLE CARD: quality close enough to shock a sceptic, subscription economics several times worse. CLUSTER: u/xuwenhao (3, the plan-value inversion: '$200 to get $10,000 API quota' on Claude vs '$250-$400 to get $3000~$4000' on Kimi), u/LargeLanguageModelo (21, '$39 plan, hit the 5hr window on 3 prompts'), u/jurapiotr (2), u/Ok_Combination4949 (6, '2.5 prompts'), u/Neither_Profession77 (1, 'token hungry (2-3x) atleast from fable'), u/monoheroshiraf (1, weekly limit in 3 days). COUNTERWEIGHTS CARRIED DELIBERATELY: u/Human_Criticism_2872 (19) finds $100 'very generous' — the complaint is TIER-DEPENDENT, not absolute; u/ThoughtHistorical596 (1) reports 'setting reasoning to Low and context window to 256k and on the $100 you get extreme usage out of it and the performance is still fantastic' — which COLLIDES with Moonshot's own doc (thinking off routes K3 to K2.6) and with its launch `reasoning_effort: max` guidance, a collision the card SHOWS rather than resolves; u/Due_Bluejay_5101 (11) warns Anthropic/OpenAI subscriptions 'have been heavily nerfed over the past 2 months', so any comparison ages fast — a caution that cuts against this card too. ⚠️ THREE COMMENTERS EXCLUDED FOR STAKE, logged so the exclusion is auditable: u/reaznval (Kimi beta tester, given free subscriptions — cannot report on paying for quota), u/gargetisha (posts for Cline), u/Prestigious_Sale_529 (table explicitly LLM-generated: 'I asked gpt-5.6-pro about this' — a model's opinion about a model is not a voice). 11 voices counted. | ROUND 5: scores re-read live 2026-07-20 and several had drifted (thread 185->183, Human_Criticism_2872 17->19, djdante 6->7, Due_Bluejay_5101 10->11, Ok_Combination4949 5->6). ⚠️ The round-5 brief supplied 187 and 20; the LIVE SOURCE read 183 and 19. The source was followed. Read all scores as a timestamp, not a fact. 🕐 DATING CAVEAT ADDED ROUND 5 (2026-07-20): djdante's '$20 Kimi plan and a $20 Claude plan → Kimi k3 will consume 4X more of your weekly quota than Fable 5 High' and ahnyudingslover's 'fable5 uses about 30-50% of my 5 hour usage limit on the 20usd plan' were BOTH measured against a $20 Claude Pro plan that INCLUDED Fable 5. From 2026-07-20 Pro no longer bundles Fable 5 (Max $100+ / Team Premium only, at 50% of limits). The measurements are valid as made and are NOT deleted; they are date-stamped on the card face as pre-07-20. Reproducing djdante's comparison today requires a $100 Max plan on the Claude side.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5b66dc000000000e036b1b?xsec_token=ABDgkwAofr4XdFCpQXCtNMc_weiknpAu7MBkkXRLgBROg=&xsec_source= 流落人间的小丑 — 241 likes / 106 comments. The BEST-MEASURED cost note in the 小红书 corpus: 「我昨天测的账:K3总花费0.71美元,Grok是0.30美元,2.4倍。15.8万token里81%是推理,代码只占19%。每行代码成本0.031美元,是Grok的3倍。75分钟跑完,Grok只要5分钟🐢」 and 「API定价K3是$15/M output,和Claude Sonnet 5完全同价。」 ⚠️ THE COMPARATOR IS GROK, NOT CLAUDE — the 2.4x and 3x-per-line figures DO NOT TRANSFER to a Claude comparison and the card labels them as Grok figures. What DOES land: (a) the 81% reasoning / 19% code token split, the cleanest published measurement of the very mechanism Flynt names independently on 掘金; (b) his Sonnet-5 price identity, which independently reproduces the card's own price-sheet finding. ⚠️ NEUTRALITY DOCKED to 0.6: the note's FRAMING is self-promotional. The numbers are first-hand; the narrative is not evidence. Counted as one voice. DURABLE KEY: note id 6a5b66dc000000000e036b1b + author + the substring 「每行代码成本0.031美元」. ⚠️ The url above is the canonical unsigned form for reference ONLY — it is REJECTED on fetch; re-read requires searching and following the SIGNED (?xsec_token=…) url.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5e0d19000000001002ae46?xsec_token=ABbmFhjLI2a_bslgp4785vwxh9jQbJi9GPt7dIvERWzeg=&xsec_source= 自由之路 — a Chinese practitioner independently stating THIS CARD'S OWN THESIS as a decision rule, which is far stronger corroboration than agreement with a conclusion: 「1. 有条件用 fable5 的,不会考虑用 Kimi,平替说法不妥 2. ... 如果你需要编码,还需要审美,kimi 几乎是少有的良好补充 3. 如果没条件用上面 2 个,且又不在乎性价比,那 kimi 是个好选择」 — (1) anyone able to use Fable 5 will not be considering Kimi, and calling it a drop-in substitute is not right = TRIAL, DON'T MIGRATE; (2) if you need coding AND aesthetic judgement, Kimi is one of the few good COMPLEMENTS = second seat; (3) if you can't get either of the above, Kimi is a good choice = the availability argument. He reached all three independently, in Chinese, from his own use. DURABLE KEY: note id 6a5e0d19000000001002ae46 + author 自由之路 + the substring 「平替说法不妥」.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5d00c1000000000f01d32a?xsec_token=ABg5BdPlstMxSPeFWIubk_Ode65Dmbq-kX_dUNiL-OOL0=&xsec_source= 小杯冰拿铁 — the THIRD independent first-hand contradiction of Moonshot's own 'excessive proactiveness' limitation, and the most interesting one because it documents the OPPOSITE failure mode: K3 RETRACTING its own conclusion. 「它之前下结论说"定时任务类扩展是生态空白",等数据拉完,自己跑回来纠正——"不对,有 57 个相关包,我之前的判断错了"」 ('earlier it concluded scheduled-task extensions are an ecosystem gap; once the data finished pulling it came back on its own and corrected itself — no, there are 57 relevant packages, my earlier judgement was wrong'). Corroborations of 'excessive proactiveness' across both sweeps and all languages: ZERO. First-hand contradictions: THREE. DURABLE KEY: note id 6a5d00c1000000000f01d32a + author 小杯冰拿铁 + the substring 「定时任务类扩展是生态空白」.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5a3df20000000022018d93?xsec_token=ABOF6h4dhCuHuZA1ibrm5Jk9RqvYK1DpncJYXCBcPlDSw=&xsec_source= Rockey — carried for TWO reasons. (1) HE CONTESTS THE ¥699 FIGURE: the card states, from V2EX @SiWXie, that the Kimi CLIENT needs ¥699/month for K3's 1M context. Rockey, same era, puts it at ¥199. Two first-hand users, same week, ¥699 vs ¥199 for the same thing — so the card now CARRIES THE DISAGREEMENT rather than either number. What both agree on, and what the card therefore does state, is the Kimi CODE tiering: ¥199 = 1M / ¥99 = 256k / ¥49 = no K3. (2) A DIRECTION data point: 「最近 GPT、Claude 疯狂重置额度…真的更想直接去用它们」 ('lately GPT and Claude have been resetting allowances like crazy… honestly it makes me want to just go use them instead') — i.e. drifting TOWARD Claude/GPT, not away. DURABLE KEY: note id 6a5a3df20000000022018d93 + author Rockey + the substring 「最近 GPT、Claude 疯狂重置额度」.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a59cd8c0000000004028637?xsec_token=AB0LmIDPbnCUqHZi9HPLl86oYR6KrxoLt3ONlmj1Ju9dk=&xsec_source= 达布卡 — 193 likes / 278 comments. READ AS A TRUST SIGNAL, NOT AS EVIDENCE ABOUT THE MODEL, and NOT counted as a sentiment voice. The card already asserted that 'several independent threads allege coordinated pro-K3 posting' with only English pseudonyms behind it; this gives the allegation a named first-party Chinese source: 「Kimi K3凌晨上线,随后就是铺天盖地的商单,都快把真实用户的声音淹没完了」 ('K3 went live in the small hours and straight after came a blanket wave of PAID PLACEMENTS — they've all but drowned out the voices of actual users'). Independently seconded by 慢悠悠记录生活: 「我知道K3刷榜很好,也请了不少KOL来宣传」 ('K3 games the leaderboards well, and a good many KOLs were hired to promote it'). This is why the card weights GitHub bug reports and same-harness numbers above sentiment volume, and says so.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a59297f0000000006011082?xsec_token=AB0LmIDPbnCUqHZi9HPLl86omjcQvFET-V4tS_3YdbDxU=&xsec_source= ⚠️ VENDOR — DISCOUNTED OUTRIGHT, QUOTED NOWHERE, COUNTED AS NOTHING. 小红书 Kimi智能助手 is MOONSHOT'S OWN ACCOUNT (uid 657fb438000000001902dbcf), and its K3 launch post — 「Meet Kimi K3!2.8万亿参数,全新底层架构」, note id 6a59297f0000000006011082, 2026-07-17 — is the HIGHEST-ENGAGEMENT K3 item in the entire 小红书 corpus. ⚠️ URL REPAIRED 2026-07-20: this row previously pointed at `/user/profile/kimi-zhinengzhushou`, which is a SLUG AND NOT A 小红书 USER ID (those are 24-hex) and RETURNED HTTP 404 — a dead link published in the ledger, the same defect class that was removed from xhs-manyouyou one row over and not applied here. It now points at the LAUNCH POST ITSELF via a SIGNED url, so the engagement claim below is checkable rather than merely asserted. The durable account handle is uid 657fb438000000001902dbcf; note that the bare profile path for it is NOT a clean link either — it answers 302 to a login interstitial and only returns 200 once the redirect is followed, which is why this row points at the signed NOTE and not at the profile. ⚠️ THE LIKE COUNT MOVED AND IS CORRECTED: 8,016 → 8,022 read 2026-07-20; the note keeps accruing and any exact figure here is a timestamp, not a constant. Listed for one reason only, and it is a reason readers should have: the single loudest voice on the platform where this card found its sharpest cost evidence is the vendor's. Iron rule #1. This is also the concrete reason engagement volume is not used as a signal anywhere on this card. ⚠️ Note that the vendor's own post concedes the frontier gap in its own words (「虽然 Kimi K3 的整体表现仍落后于最强的闭源模型 Claude Fable 5 和 GPT-5.6 Sol」) — read, and still carrying nothing, because a vendor claim in either direction is not evidence here.
- youtube — www.youtube.com/watch?v=uuxJKsN4O88 Creator Magic — 22,148 views. PAID THE $19 HIMSELF, NO SPONSOR, which is why it is here at all when two higher-reach K3 videos are disqualified. Unlike the B站 videos this one HAS A TRANSCRIPT, so spoken content is citable. Verdict: 'out of three of the four tasks I believe Fable five was in the lead and faster and included in my plan. And I didn't hit API errors… I'm still going to be pumping my money into Anthropic and Fable five.' Availability: 'I've hit a 403 on all of them. I've hit my usage limit literally in just a few minutes with $19'. ⚠️ AND THE METHODOLOGICAL REBUTTAL TRAVELS WITH IT — top comment @tutkarz (114 likes): 'so you compared $200 Fable to $19 Kimi K3. Fair comparison, indeed.' HE NEVER STATES HIS CLAUDE TIER, and the card says so. That single unstated variable disqualifies the video as a MEASUREMENT while leaving it perfectly good as a first-hand report of hitting a $19 wall in minutes — which is the only thing it is used for. @tutkarz is quoted but NOT counted as a sentiment voice: he is judging the comparison, not the model.
- youtube — youtube: AICodeKing — Kimi K3 review (description reads "Sponsored by Moonshot AI") (search query, not a page) ⚠️ DISQUALIFIED ON IRON RULE #1 — VENDOR-FUNDED. The description discloses 'Sponsored by Moonshot AI'. Quoted nowhere, counted as nothing. Listed because a disqualification a reader cannot see is indistinguishable from a source we never found — and this is one of the highest-reach K3 'tests' in English.
- youtube — youtube: Dubibubi — Kimi K3 review (affiliate link, "get 15% bonus usage") (search query, not a page) ⚠️ DISQUALIFIED ON IRON RULE #1 — AFFILIATE. Carries a discount/referral offer ('get 15% bonus usage') for the product under review. Quoted nowhere, counted as nothing. Listed for the same reason as AICodeKing: the exclusions have to be visible to be worth anything.
- v2ex (CN) — www.v2ex.com/t/1227856 Thread 'KIMI 的 K3 上线了,但开会员你也用不了满血的' — OP @SiWXie, 61 replies, re-read line-by-line 2026-07-20. The OP's own screenshot-backed post is the REAL source for the tier->context table (¥49 no K3 / ¥99 256k / ¥199+ 1M Kimi Code / ¥699 client 1M) — the card previously mis-cited that to Moonshot's doc page. Carries @ProphetN #57 (whose words were FABRICATED in the previous draft — see quotes.md §Z), @operapeking #55 (the steelman: Codex subscriptions are 256k-capped too), @mon6912640 #41 and #61, @hm279 #48, @Msxx #53 (astroturf flag), @cat9life #18, @qf19910623 #34, @tarikzhang #47. NOTE: @SiWXie is OP *here*, but his token measurement is in t/1228031 where he is a replier — the previous draft merged the two.
- v2ex (CN) — www.v2ex.com/t/1228031 Thread 'kimi k3 买的 199 会员,终于有随便用的感觉了' — OP @ujujzhaos, 36 replies, re-read line-by-line 2026-07-20. Note the thread's framing is POSITIVE (OP says the ¥199 tier finally feels unlimited) and the replies largely contradict him — useful, because it is not a pile-on thread. CORRECTION: the card previously cited this thread for @m1nm13's '199 ≈ claude pro fable allowance' quote. @m1nm13 DOES NOT APPEAR IN THIS THREAD; that quote is reply #40 of t/1227894 and the href has been fixed. What this thread genuinely carries: @SiWXie #29+#30 (the only first-party token measurement: ~32M/5h, 164M/week, 660M/month, scoped to ¥199 running K3 throughout, measured via ZCode), @a566 #35 (1.7MB / 60-file read = 80% of the 5-hour quota), @whyso #20, @poxiaohy #33 (effort tiers appearing mid-sweep).
- v2ex (CN) — www.v2ex.com/t/1227894 Thread 'Kimi K3,不来个朴实无华的讨论吗?' — OP @lelelelelelele, 59 replies, re-read line-by-line 2026-07-20. THE thread for this card's cost evidence. Carries: @m1nm13 #40 (¥199 K3 allowance ≈ 'claude pro flabe' allowance — source's own typo, reproduced [sic]); @106npo #32 (the 4x-vs-10x plan-multiplier argument, hedged by his own 我记得/'I recall', so the card does not carry it as a number) and #23 ('价格已经和 sonnet 一致了' — independently CONFIRMED against Anthropic's price sheet); @dingawm #19 (hostile corroboration of the frontend praise); @yvescheung #47; @whyso #57; @fcten #48 (routing, not switching); @owen800q #2 relays the Frontend Code Arena table. Four quotes were attributed to the wrong thread in the previous draft and are corrected. 🕐 DATING CAVEAT ADDED ROUND 5: @m1nm13's comparator — 'a Claude Pro Fable allowance' — ceased to exist on 2026-07-20, when Pro stopped bundling Fable 5. The ¥199 reading stands as measured 2026-07-17.
- v2ex (CN) — www.v2ex.com/t/1228182 ⚠️ CLAIM WITHDRAWN 2026-07-20. This source previously carried 'receives no substantive backend answer in 8 replies' — an evidence GAP. Re-read IN FULL on 2026-07-20: the thread has 13 replies, and reply #2 (@thinkeryu, 2026-07-18 12:23:20 +08:00 — i.e. already present when the sweep read it, so the claim was wrong when written, not merely stale) is a substantive first-hand BACKEND report and a POSITIVE one: 「我在自己的 rust 项目上用,除了慢一点,完成度很高,没有低级错误,比较靠谱」. What the source now supports is the weaker, true statement: backend use is barely measured (n=1) and that n=1 is favourable. The OP's framing — every K3 test in circulation is frontend — still stands and is separately quantified by the Qiita 101-use-case catalogue (78.2% game/frontend/3D).
-
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a59b6b3000000000f01609f?xsec_token=AB0LmIDPbnCUqHZi9HPLl86pe2aYik96FNXDNQZWjFwVk=&xsec_source=
Directly tests Moonshot's own '>90% cache hit keeps cost low' defence: user verified a 94% hit rate and still exhausted the 5-hour window in two tasks, concluding it would be dearer than Claude Fable 5. A vendor claim falsified by a user's own console. ✅ RE-VERIFIED IN FULL 2026-07-20 VIA agent-reach — the 'NOT RE-VERIFIABLE / login-walled' flag this entry carried was FALSE (bug #39). 小红书 requires a SIGNED url (?xsec_token=... returned by `search`); the unsigned /search_result/
that rounds 1-3 fetched is rejected and returns a JS shell, which was misread as a login wall. ⚠️ COMMENT COUNT CORRECTED 59 → 75 (the note kept accruing comments after the original capture). DURABLE RE-VERIFICATION KEY (signed urls expire, so do not rely on one): note id 6a59b6b3000000000f01609f + author 幸运的蜗牛 + the body substring 「命中率都是 94%」. - xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5b967a000000000f008253?xsec_token=ABDgkwAofr4XdFCpQXCtNMcz5GjhjPSYoJXpIQ2O3H2HI=&xsec_source= Screenshot-backed: rate-limited fifteen minutes after paying for the ¥99 plan, at 18.5% of the 1M context, during a first code review. Author sought a refund. Screenshots are why this scores high despite being a single angry post. ✅ RE-VERIFIED IN FULL 2026-07-20 VIA agent-reach — the 'NOT RE-VERIFIABLE / login-walled' flag was FALSE (bug #39, see round4_note). ⚠️ MISATTRIBUTION CORRECTED: this author was labelled on the card as 'V2EX @blueraincoat / XHS'. THERE IS NO V2EX EVIDENCE FOR THIS PERSON — the V2EX half was invented by association. He is a 小红书 author and nothing else; the card now reads 'XHS blueraincoat'. Full note title: 「KIMI K3? 什么情况?吃相有点难看吧?」. ⚠️ COMMENT COUNT CORRECTED 36 → 49. He also supplies a DIRECTION data point: he arrived from Codex/GLM/DeepSeek, NOT from Claude. DURABLE KEY: note id 6a5b967a000000000f008253 + author blueraincoat + the substring 「1M的上下文才跑到18.5%就限额了」.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5b979b000000000c014b1b?xsec_token=ABDgkwAofr4XdFCpQXCtNMc_xPUmIdEFVmPId4lxeaYyI=&xsec_source= The best vision-in-the-loop counter-evidence found: a trivial annotated-screenshot padding edit that 'any GPT or Claude model' handles, where K3 spent ages treating width as height. Same author wired K3 INTO Claude Code because Kimi's own CLI/plugin cut out mid-task — a routing outcome, not a switch. ✅ RE-VERIFIED IN FULL 2026-07-20 VIA agent-reach — the 'NOT RE-VERIFIABLE / login-walled' flag was FALSE (bug #39, see round4_note). ✅ ALSO SUPPLIES A DIRECTION DATA POINT, new this round: 「2026年7月,推荐的ai coding配置依旧是100刀gpt pro, 20刀claude,gpt当主力」 — his July 2026 recommendation is $100 GPT Pro + $20 Claude with GPT as the workhorse, i.e. he is not substituting AWAY from Claude. DURABLE KEY: note id 6a5b979b000000000c014b1b + author momo + the substring 「但是 k3 能把宽度认为是高度调半天」.
- xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5a1afa000000000503b41c?xsec_token=ABOF6h4dhCuHuZA1ibrm5JkyPasAKhz-AzP7olAgcZ08w=&xsec_source= This is XHS 终梦. Recomputed 7 days / 4000+ requests / 850M real tokens at K3 list pricing. ⚠️ COMPARATOR IS DEEPSEEK, NOT CLAUDE — the 48x multiple does NOT transfer and the card does not use it. What transfers is the MECHANISM: agent loops are cache-read dominated, so cache-read price drives real cost, invisible in headline $/MTok tables. ✅ RE-VERIFIED IN FULL 2026-07-20 VIA agent-reach — the 'NOT RE-VERIFIABLE / login-walled' flag was FALSE (bug #39). 48 comments / 40 likes. ✅ RECOVERED THE MISSING MECHANISM NUMBER the card had been asserting without a figure: 「DeepSeek ¥0.025/百万token,约等于白送 K3 ¥2/百万token,整整 80 倍」 and 「输出价更离谱,¥6 vs ¥100」 (both DeepSeek-vs-K3, NOT Claude). 🚨 FABRICATION-CLASS DEFECT FOUND HERE AND FIXED: his conclusion was printed on the card and in quotes.md INSIDE 「」 UNDER A HEADER READING 'Verbatim' as 「日常杂活 DeepSeek,最难的活给 K3」. That is a TRUNCATION, not what he wrote. Real text: 「我的方案:日常杂活 DeepSeek 焊死,最难的活给 K3 单独开小灶。」 The reading (routing, not switching) survives; the quotation did not. Treated with the seriousness of the round-1 fabricated quote. DURABLE KEY: note id 6a5a1afa000000000503b41c + author 终梦 + the substring 「agent 用法 99% 的 token 都是上下文反复重读」.
- personal blog (CN) — zhiqiang.org/it/kimi-quota.html THIS IS 阅微堂. ✅ NAMING FIXED 2026-07-20: the round-4 brief reported 阅微堂 as 'UNFINDABLE — zero hits across nine varied searches, and no URL exists in the card or receipts', and proposed removing it from the coverage list. THE FILES DISAGREE AND THE FILES ARE RIGHT — 阅微堂 is simply this blog's own name, and zhiqiang.org has been in quotes.md and in this ledger with a live URL throughout. It was unfindable only because the coverage footer printed the Chinese blog name with NO DOMAIN BESIDE IT. Fixed by naming it '阅微堂 (zhiqiang.org)' in the footer, NOT by deleting a source that resolves. Highest-neutrality source in the Chinese sweep — personal academic blog, no monetisation, reproducible UI path. Documents a THIRD, hidden monthly quota tier split into separate kimi and code pools that the coding console never surfaces. Pre-K3 but structural, so it still holds.
- 51cto blog (CN) — blog.51cto.com/lizhuo6/14771976 ⚠️ RESTORED 2026-07-20 AND FLAGGED; ✅ UPGRADED IN ROUND 4 — IT NOW HAS TWO WITNESSES. Carries the card's only 'disposition, not ranking' reading — K3 as a 'strong executor' engineer rather than a 'proactively refactoring' architect — plus a 「我跑了十几个用例就烧掉了接近 30 美元」 (~$30) burn line. THE LIVE PAGE IS STILL UNREACHABLE: re-checked 2026-07-20, it returns HTTP 567 with a 7,351-byte bot-wall interstitial as its body, and agent-reach returns an EMPTY body. BUT Exa's semantic index carries the FULL article text and it MATCHES BOTH of the card's quotes VERBATIM — 「我跑了十几个用例就烧掉了接近 30 美元」 and 「K3 的编程能力更像"执行力强"的工程师,而不是"会主动重构"的架构师」 — obtained by a route with no access to our capture. That is a second, independent witness to the string. The card-face flag therefore changes from 'quoted from capture — not re-verifiable' to 'live page returns HTTP 567; text independently corroborated by a third-party index'. Reliability raised 0.4 -> 0.55 on the second witness, and no further: a corroborated string is not a readable page. Round 2 CUT this source and subtracted its voice while retaining four 小红书 notes with what was then believed to be the identical defect; that was two standards for one problem. Note that in round 4 the 小红书 'defect' turned out not to exist at all — so 51CTO is now the ONLY genuinely hard-to-check source on this card. It remains the weakest item and is labelled as such.
- note.com (JP) — note.com/masa_wunder/n/n114463d0c7f0 $40 spent across three harnesses. The most methodologically disciplined voice in any language: explicitly warns against the 'beats Opus / beats Fable' framing, argues harness mesh matters more than raw model strength, and self-flags his own n=1 on throughput. Third card running, a note.com post is among the most neutral things found.
- tech press (JP) — pc.watch.impress.co.jp/docs/news/2125917.html The clearest statement anywhere that the K3 LICENCE IS UNANNOUNCED — the official blog never says what it will be — with K2.6/K2.7's Modified MIT given as precedent only. Established trade press, no stake in either vendor.
- qiita (JP) — qiita.com/yuma_morita/items/6489e1acc5c31192bbea Catalogued 101 public K3 use cases and found 78.2% are game / frontend / 3D. Quantifies what the V2EX backend thread noticed anecdotally. Neutrality docked — it is a hype-adjacent roundup — but the distribution itself is checkable and it cuts against the model's breadth claim.
- consultancy — www.nxcode.io/resources/news/kimi-k3-benchmarks-coding-agent-evaluation-guide-2026 Sells neither model. Articulates the methodological spine of this card better than any other single source: 'the harness is part of the product', and there is not yet enough same-harness independently reproduced evidence for a universal best-coding-model claim. Neutrality docked — a consultancy has an interest in 'you need an evaluation'.
- vendor docs — www.kimi.com/code/docs/third-party-tools/claude-code.html ⚠️ VENDOR — primary for its own product's behaviour. ✅ URL CORRECTED 2026-07-20: the previously cited path …/third-party-tools/other-coding-agents.html now returns 404, which is how a 2026-07-10 snapshot (six days BEFORE K3 launched) was read as live. The working page is …/third-party-tools/claude-code.html (HTTP 200, re-fetched 2026-07-20). ✅ THE ROUND-1 WITHDRAWAL OF THIS QUOTE WAS ITSELF AN ERROR: 「关闭 thinking 后 K3 和 K2.7 Code 都会被路由到 K2.6,请保持 thinking 开启以使用 K3 / K2.7」 IS on the live page, verbatim, in a 注意 callout — turn thinking off and you are served K2.6, one keystroke away (⌥T / Alt+T). Only the 'silently, without being told' gloss was wrong: it is explicitly documented. Two further claims the old draft made about this page are REFUTED: 'K3 is not named on the page' (K3 appears 9 times) and 'caps you at 262144, not 1M' — the live tier table reads Andante kimi-for-coding 262144 · Moderato k3/kimi-for-coding 262144 · Allegretto 及以上 K3 用 1048576, K2.7 Code 系列用 262144. So 262144 is the Andante/Moderato cap and K3 gets the full 1M at Allegretto and above.
- vendor docs — platform.kimi.com/docs/guide/kimi-k3-quickstart VENDOR — used only for hard integration constraints: multi-turn and tool calls must echo the complete assistant message back VERBATIM; reasoning_effort supported max only at launch; vision rejects public image URLs; web-search tool 'not recommended for production'. ⚠️ Contradicted Moonshot's other doc site live during the sweep — every claim from here is dated. RE-CHECKED 2026-07-20: HTTP 200. The verifier's earlier '000' was transient, not a dead link — the entry stands.
- vendor docs — platform.kimi.ai/docs/guide/claude-code-kimi VENDOR — the Claude Code integration page, hence directly on-topic. Requires ENABLE_TOOL_SEARCH=false because the Kimi endpoint does not support Claude Code's Tool Search or 'tool calls misbehave'. Reliability docked slightly: this site and platform.kimi.com disagreed with each other during the sweep.
- industry media (CN) — www.woshipm.com/ai/6431989.html The ONLY source found documenting the reverse migration direction — developers whose task didn't run through and who switched back to Claude Opus. Reliability capped because it is second-hand aggregation with no named developer. Reported on the card as thin, not as a finding.
- serp — google: "Kimi K3 vs Claude" (search query, not a page) Page one: artificialanalysis.ai, llm-stats.com, incrypted.com, finance.biggo.com, futurism.com, tomshardware.com, gizmodo.com, myclaw.ai, layer3labs.io, kie.ai. FOUR of ten are commercial AI properties selling access or leads; two more are SEO aggregators; three are wire rewrites. ZERO independent practitioner reviews.
- serp — google: "Kimi K3 vs Claude Fable 5" (search query, not a page) Page one: codersera.com, aitoolsreview.co.uk, vallettasoftware.com, digitalapplied.com, llm-stats.com x2, medium.com, benchlm.ai, artificialanalysis.ai x2. SIX of ten SEO/affiliate/agency-owned. Recorded as the meta-signal, not read for verdicts.
- reseller / SEO — kie.ai/blog/kimi-k3-vs-claude READ, DELIBERATELY NOT QUOTED. Representative of the page-one cluster (kie.ai, myclaw.ai, layer3labs.io, llm-stats.com, codersera.com, vallettasoftware.com, digitalapplied.com, kimi-k2.org, ofox.ai). These sell API access, leads or agency services. Checked for cross-checkable facts only; they recycle vendor benchmark tables — including the KimiCode-harness 88.3 — with no methodology note. Never carries a verdict here.
- hacker news — news.ycombinator.com/item?id=48935342 'Kimi K3: Open Frontier Intelligence' — ⚠️ COUNT CORRECTED 2026-07-20 via the HN Algolia API: 1,206 comments / 2,090 points. This entry and the card said '~1,850 comments, swept in full', and NO SUCH THREAD EXISTS: 1,206 + 598 (item 48960218) = 1,804 — TWO SEPARATE THREADS HAD BEEN SILENTLY ADDED TOGETHER AND PRESENTED AS ONE. The three real threads are now listed separately with their real counts. Source of the 86-token 'hi' finding. Also used as a NEGATIVE check: the three threads were grepped for proactive/improvis/overeager to test Moonshot's own 'excessive proactiveness' admission and got ZERO hits — which is why that limitation stays vendor-disclosed rather than corroborated.
- hacker news — news.ycombinator.com/item?id=48960218 'The Kimi K3 Moment' — 598 comments / 623 points (HN Algolia API, 2026-07-20). The HN discussion of the Bochinski 'I can't tell them apart' post. ADDED IN ROUND 4: this thread was previously invisible on the card because its comment count had been merged into the main thread's to produce a fictitious '~1,850-comment thread'. Listed separately so the corpus the card claims to have swept is the corpus it actually swept.
- hacker news — news.ycombinator.com/item?id=48947717 The HN thread on simonwillison.net's K3 writeup — 221 comments / 403 points (HN Algolia API, 2026-07-20). ADDED IN ROUND 4 for the same reason as hn-k3-moment: the card's HN claim named one thread and counted two. Three threads, 1,206 + 598 + 221, all now named with real counts.
- reddit — reddit search ×5 phrasings (authenticated session): "Kimi K3 did things I didn't ask / went rogue / changed files / unexpected decisions / overeager" (search query, not a page) NO ON-TOPIC RESULT. ⚠️ WORDING CORRECTED 2026-07-20: this was logged as 'an empty result set', which is misleading under the round-4 tooling. A non-matching SEMANTIC search does not come back empty — it comes back with unrelated high-scoring posts, which looks nothing like emptiness, and describing it as 'empty' overstates how clean the null is. What actually happened, stated precisely: FIVE separate phrasings were run through an AUTHENTICATED Reddit session on 2026-07-20 and none returned an on-topic result. Logged because a null result is evidence: it is the second independent check that fails to corroborate Moonshot's 'excessive proactiveness' limitation — which now also has THREE first-hand contradictions against it (XHS Z3PH1NUE ⚠️comparator GPT-5.5 not Claude, the ShipSolo writeup, and XHS 小杯冰拿铁 documenting K3 RETRACTING its own conclusion) and ZERO corroborations.
- multi-platform — sweep: K3 self-hosting reports, all languages (search query, not a page) RETURNED NOTHING. Zero observed self-hosters of K3 in any language. Everyone testing it is on the Kimi API, the Kimi subscription, OpenRouter or opencode-go. Every self-hosting discussion found is hypothetical or recycled from K2-era guides. Directly limits what the 07-27 weights promise is worth on day one.
- multi-platform — sweep: controlled Kimi Code vs Claude Code / Cursor quality delta (search query, not a page) STILL THE MISSING NUMBER — but ⚠️ NARROWED 2026-07-20, NOT DELETED. No independent test has run THE SAME K3 TASKS in both Kimi Code and a third-party harness, so the harness penalty itself remains unquantified. However, 'nobody has measured' overstated the silence: K3 HAS now been run publicly INSIDE CLAUDE CODE (AI超元域, B站, 2026-07-17) and inside a shared agent harness alongside rivals (Token就是词元, B站, 2026-07-18). All three candidate videos were read (agent-reach first; Chrome MCP once, operator-authorised, as the sanctioned last resort). WHY EACH FAILS AS A HARNESS MEASUREMENT: Creator Magic ran K3-in-Kimi-Code vs Fable-5-in-Claude-Code — TWO VARIABLES MOVED AT ONCE, plus an undisclosed Claude tier against a $19 Moonshot plan; AI超元域 ran K3 in Claude Code but its Kimi-side tasks were DIFFERENT TASKS, so there is no common workload to difference; Token就是词元 ran multiple models in ONE shared harness, which isolates the MODEL, not the harness — the opposite of the needed experiment. Also relevant to why this could not simply be read out of the videos: BOTH B站 videos have NO SUBTITLE TRACKS — /x/player/v2 returns `code 0` OK with `"subtitles": []` for cid 40037188863 and cid 40067599223. That is a definitive negative, not a fetch failure.
- zenn / hatena (JP) — sweep: zenn.dev + はてなブックマーク "Kimi K3" (search query, not a page) RETURNED NOTHING substantive. Zenn and Hatena had no hands-on K3 coding writeups at sweep time. Listed so the Japanese coverage claim is not overstated — the JP signal came from note.com, Qiita and PC Watch only.
- korean dev web — sweep: Velog + tistory "Kimi K3" (search query, not a page) RETURNED NOTHING substantive — no first-hand Korean K3 coding reports at sweep time. Absence is reported, not assumed.
- wechat 公众号 (CN) — weixin.sogou.com search: "Kimi K3 实测" (search query, not a page) NOT SWEPT — Sogou's WeChat search was inaccessible during the sweep. A real gap, disclosed rather than papered over: 公众号 is vendor-heavy and would mostly have been discounted anyway, but that is a prediction, not a finding.
- bilibili (CN) — www.bilibili.com/video/BV1CDNd65EDc ❌ THIS ENTRY'S ORIGINAL NOTE WAS A CHARACTERISATION OF MATERIAL NOBODY HAD READ, and it is withdrawn. It said: 'Swept, read, NOT quoted. Chinese K3 video reviews at launch week were demo-and-reaction content — frontend showcases with no methodology, no harness disclosure and heavy promotional framing.' B站 WAS NEVER SWEPT. Rounds 1-3 never called agent-reach, which reaches B站 fine. Asserting what a whole platform contains without reading it is an iron-rule-#2 violation independent of bug #39, and it happens to have been FALSE: the round-4 sweep found 13 independent uploaders with first-hand K3 material, several of whom DO disclose their harness. This entry is retained at low reliability as the record of the error; the real B站 sources are the bili-* entries below.
- bilibili (CN) — www.bilibili.com/video/BV1GRKJ6fEgn AI超元域 — the most decision-relevant B站 find, and the reason the card's 'nobody has run K3 in a Claude-shaped harness' silence is narrowed. ⚠️ STRICT CITATION LIMIT, ENFORCED: this video has NO SUBTITLE TRACK. /x/player/v2 returns `code 0` (OK) with `"subtitles": []` for cid 40037188863 — a DEFINITIVE NEGATIVE, not a fetch failure. Therefore only its EXISTENCE, the uploader's own WRITTEN description/replies, and COMMENTS are cited; spoken content is never quoted or characterised. ❌ ROUND 5 — ONE UPLOADER QUOTE CUT AS UNCONFIRMABLE. 「三个小时以上 中间因为额度达到限制 三个小时后才能继续开发,所以耗时了将近4小时」 WAS CARRIED AS UPLOADER-WRITTEN AND VERBATIM AND COULD NOT BE FOUND: not in the full description, not in 184 distinct comments and sub-replies (including ALL EIGHT of the uploader's own replies), not in reply-API pages 1-3 (60 replies + subreplies). Its COMPANION quote matched exactly by the same route, so the route works — the string simply is not where the card said it was. DELETED everywhere, together with the 'nearly four hours' claim resting on it. An unconfirmable quote is treated as FABRICATED until proven otherwise; that is the round-1 lesson. ✅ CONFIRMED VERBATIM and retained (uploader replying to @xzo_ozc's 「定价飘了」): 「定价确实贵 199订阅演示完视频中的案例消耗了周额度的20%多」 (on ¥199, demoing this video's cases consumed >20% of the WEEKLY quota). Top independent comment @lee老头儿的幸福生活 (181 likes) on over-thinking and 'chicken-rib' pricing; and @捌玖柒897 reports the §G footgun in the wild — ccswitch into Claude Code, model set to k3, k2.6 actually called — UNANSWERED. ✅ ROUND 5 ADDITIONS, both CONFIRMED VERBATIM via the public reply API: @若迹铭 「这么多测试才用了21%的周额度吗?我也是同等套餐,挑的非忙碌时段,手机app上做个后室用了11%月额度,看来还是得kimi code。」 and the uploader's reply 「主要是不刷新额度 要是学codex刷新额度就好了」 ('the main thing is the quota doesn't refresh — it'd be good if they copied Codex and refreshed it') — the uploader, who has every incentive to be positive about the model he is demoing, naming the STRUCTURAL complaint rather than the price. And @DD蛋卷 「实际上写后端还是很差,写前端真的强无敌,比gpt5.6sol好太多…不过kimi额度消耗是真夸张,99套餐一晚上用了1/3,还是太贵了」 — the frontend/backend split, first-hand, landing on both sides at once. ALL ATTRIBUTED AS COMMENTS, never as spoken content (both videos return "subtitles": []). ⚠️ COMMENT-SET DISCLOSURE, RESTATED IN ROUND 5 because the old one UNDER-reported the read: the full set is login-walled (「登录后查看 372 条评论」) and requires WBI signing. What was actually read: the PUBLIC REPLY API, PAGED, HOT-SORTED — pages 1-3, 60 top-level replies plus sub-replies; an adversarial verifier independently pulled 184 distinct comments by the same route. So 107 of 372 replies (pages 1-3, ~29%) are what this card's method statement rests on; an adversarial verifier separately pulled 184 by the same route, and the card claims the LOWER, directly-reproduced figure. The rest was not read. An earlier draft said 'roughly half', which overstated the read by ~1.7x and was corrected 2026-07-20. Hot-sorting biases the first pages toward high engagement, so this is not a random sample either. (The previous 'the 20 comments read' understated the sweep AND implied a completeness ceiling never tested — an inaccurate method statement is inaccurate in either direction.) ⚠️ PROVENANCE: the BV id is as recorded by the agent-reach sweep; the durable key is uploader + date + the quoted reply text.
- bilibili (CN) — bilibili: Token就是词元 — K3 in a shared multi-model agent harness, 2026-07-18 (no BV id recorded by the sweep) (search query, not a page) CITED FOR EXISTENCE ONLY, and deliberately given no URL because the sweep did not record a BV id and this card does not assert URLs it cannot show. It matters for exactly one reason: together with AI超元域 it NARROWS the card's 'nobody has measured the harness penalty' claim. It does NOT close the gap — running multiple models in ONE shared harness isolates the MODEL, not the harness, which is the opposite of the experiment needed. No quotation is taken from it. ⚠️ Also no subtitle track: /x/player/v2 returns `code 0` with `"subtitles": []` for cid 40067599223.
- bilibili (CN) — www.bilibili.com/video/BV1X4Kz6UEZd 熊本的AI百宝箱 — uploader-WRITTEN description, negative and specific: 「1)慢。思维链太重了…2)贵。常规开发场景,199档位的5h限额基本只够用1h。」 ('1) Slow, the chain of thought is far too heavy. 2) Expensive — in ordinary development the ¥199 tier's 5-hour cap is basically only good for about an hour'). Independently reproduces BOTH halves of the card's cost finding, in Chinese, on the ¥199 tier. Description text only — no spoken content cited.
- bilibili (CN) — www.bilibili.com/video/BV1DmKw66ExG 作业茶壶君 — uploader-written: 「k3这个模型能力上毋庸置疑,但是性价比和codex的订阅相比还是拉胯」 ('K3's raw ability is beyond doubt, but on value for money it's still limp compared with a Codex subscription'). ⚠️ COMPARATOR IS CODEX, NOT CLAUDE — counted as a voice, not used as Claude evidence. Description text only.
- bilibili (CN) — www.bilibili.com/video/BV1UrKn64ECZ kate人不错 — uploader-written: 「综合编码能力排第三,仅次于 Fable 5、GPT-5.6 Sol」 ('third on overall coding ability, behind only Fable 5 and GPT-5.6 Sol'). Also the venue for the B站 community verdict on the Java-backend test, @TwinkYou (69 likes): 「和预想的差不多,远强于4.8,稍逊于fable5」 ('far stronger than 4.8, slightly behind fable5') — a lay Chinese comment landing unprompted on exactly the shape two independent eval houses measured: K3 BETWEEN the two Claudes. Description + comments only.
-
zhihu (CN) — www.zhihu.com/question/2061204677446964906
⚠️ RELABELLED 2026-07-20 — the old description ('swept, returned nothing') was wrong in BOTH halves. What is actually true: 知乎 ARTICLE PAGES READ FINE (zhuanlan.zhihu.com/p/
is fetchable via agent-reach), but 知乎 SEARCH IS CAPTCHA-WALLED — zhihu.com/search returns, verbatim, 「系统监测到您的网络环境存在异常」 and 'This page maybe requiring CAPTCHA' [sic, the site's own broken English]. THEREFORE ENUMERATION WAS NEVER ACHIEVED: we cannot say what 知乎 contains, only what the pages we could reach contain. That is a real, quoted gap and it is a NARROWER claim than 'returned nothing' — which asserted knowledge we did not have. - zhihu (CN) — zhuanlan.zhihu.com/p/2061380200647205373 ⚠️ VENDOR. The one substantive 知乎 item the round-4 sweep could reach — and it is Moonshot's OWN launch blog, reposted. It cannot carry a verdict, it is not quoted, and it is not counted as a voice. Listed because the honest description of 知乎 on this card is 'article pages readable, search CAPTCHA-walled, and the one thing we could read was the vendor's own post' — which is a far more useful disclosure to a reader than 'returned nothing'.
- juejin (CN) — juejin.cn/post/7663097122267512859 Swept, read, not quoted. An integration tutorial (how to point Claude Code at the Kimi endpoint) rather than a verdict — useful confirmation that the Claude-Code-onto-K3 route is what people actually do, but it carries no independent quality judgement. ⚠️ THE CARD'S CLAIM THAT 掘金 'RETURNED NOTHING' WAS FALSE, and the cause is now known, mechanical and reproducible: 掘金's search TOKENISES THE SPACE IN 'Kimi K3' BADLY. `?query=Kimi%20K3` returns only 1-3-year-old PRE-K3 posts; `?query=K3&type=2&sort=1` returns a dense CURRENT cluster. Rounds 1-3 ran the first form, got stale results, and concluded the platform was empty — while the best 掘金 source on this topic (juejin-flynt, below) sat in the second. The trap is now recorded in docs/SWEEP-SOURCES.md so it cannot recur.
- juejin (CN) — juejin.cn/post/7663725308165701678 Flynt — THE 掘金 find, and one of the strongest new sources on the card. API hands-on, HIS OWN HARNESS, per-task console token counts, NO affiliate link. Valuable because it reaches the card's cost conclusion by a mechanism nobody else names: FORCED MAXIMUM reasoning_effort inflating billed output on tasks that do not need it. 「同一个简单任务("把按钮的 border-radius 改成 8px"),K3 产生的 reasoning token 是 Claude `thinking: low` 模式的 6 倍。简单任务的性价比,K3 其实不占优势。」 — 6x Claude's reasoning tokens on a trivial task. Also a specific ROBUSTNESS finding against Claude: 「它在清理 observer 的时候没处理组件卸载的竞态条件……这个问题 Claude 基本不会犯」 (missed an unmount race condition Claude 'basically doesn't make'). And a genuinely MIXED verdict line neither camp would write: 「K3 用 Fable 5 大约三成的价格,做到了八九成的体验」 (eight or nine tenths of the experience at about three tenths of Fable 5's price). Console figures: task 1 ≈¥1.27, task 2 ≈¥2.37, task 3 ≈¥4.49; reasoning tokens 8,340 / 15,720 / 28,500. Scores high: first-hand, API (no plan-tier confound), quantified, no commercial stake found.
- vendor pricing page — claude.com/pricing ⛔ SUPERSEDED IN PART, 2026-07-20 (round 5). ⚠️ VENDOR (Anthropic) — primary for plan prices and tier structure, and nothing else. Fetched 2026-07-20. Free $0 / Pro $20mo or $200yr / Max 5x $100 / Max 20x $200; Team $20 standard or $100 premium per seat annual; Enterprise $20/seat + usage at API rates. THE LOAD-BEARING READING: in the page's own plan-comparison table the 'Fable' row carries NO checkmark under Free, Pro, Max 5x or Max 20x, while Opus is ticked for Pro and both Max tiers and Sonnet/Haiku for all four; on the Team & Enterprise table Fable is ticked 2 of 3. METHOD DISCLOSED because it matters: the ticks are inline SVG glyphs, counted between consecutive row labels in the fetched HTML — this is a table reading of a marketing page, not a sentence Anthropic wrote, and the card states only 'Fable 5 is not presented as an included model on any individual plan'. Also the source for Anthropic's own metering shape: 'usage limits that reset on a rolling five-hour session window, and paid plans add weekly limits on top' — which is why the card refuses to present 5-hour windows as a Kimi pathology. ⛔ THE 'LOAD-BEARING READING' ABOVE IS NO LONGER THE CARD'S CLAIM. On 2026-07-20 Anthropic made Fable 5 a permanent part of Max 5x, Max 20x and Team Premium at 50% of limits (anthropic-fable-permanent-x + six independent outlets). Max IS an individual plan. The tick-table reading is retained as the record of what this page rendered when it was read — and as a caution: a marketing table read by SVG path signature is a weak instrument, it was the ONLY support for a four-place card-face claim, and it was contradicted by the vendor's own announcement two days earlier. Prices from this page (Free $0 / Pro $20 / Max 5x $100 / Max 20x $200 / Team $20 standard, $100 premium per seat) are unaffected and remain corroborated by PCWorld.
- vendor support doc — support.claude.com/en/articles/11049741-what-is-the-max-plan ⚠️ VENDOR (Anthropic). Fetched 2026-07-20. Max 5x $100/mo, Max 20x $200/mo; Max buys 'priority access to our newest features and models'; Max carries TWO weekly limits, one across all models and one for Sonnet models only. Quoted for the sentence that matters most to this card's access axis, and which is against interest: 'we may limit your usage in other ways, such as weekly and monthly caps or model and feature usage, at our discretion' — i.e. Anthropic reserves model-level gating in writing, the same lever Moonshot pulled with its ¥49/¥99/¥199 context tiers. Neither vendor's model-availability is a fixed property, so every claim on this axis carries its date.
- vendor support doc — support.claude.com/en/articles/11049762-choose-a-claude-plan ⚠️ VENDOR (Anthropic). Fetched 2026-07-20; page itself dated 2026-05-19. Cross-check on the tier table only (Free $0 / Pro $20mo-$200yr / Max 5x $100 / Max 20x $200). Read to confirm the pricing page's numbers rather than trusting one vendor page alone. Notably it describes tiers ONLY by usage capacity and says nothing about which models each tier exposes — an absence the card reports rather than fills in.
- 小红书 — www.xiaohongshu.com/search_result/6a599595000000000101cfc6?xsec_token=AB0LmIDPbnCUqHZi9HPLl86nwRx3wcI8MZw6TI-rH0zCg=&xsec_source= ⚠️ ADDED IN ROUND 5 — this voice was NAMED ON THE CARD FACE for two rounds with NO ledger row, no date and no durable key, while every other 小红书 note had one. Iron rule #2 violation, now fixed. Note id 6a599595000000000101cfc6, uid 68309b5600000000180237a6, title 「杂谈 | Kimi K3 很强,但是呢......」, 85 likes / 22 collects / 141 comments. Durable key: note id + author + 「今天正式版干活 2h 用掉月度额度 10%」. Read verbatim via agent-reach 2026-07-20 (search-then-signed-URL). Used on the card as ONE OF TWO first-hand contradictions to Moonshot's disclosed 'excessive proactiveness' limitation: 「也不乱改」. ⚠️ THE COMPARATOR IS NOT CLAUDE. The note writes only 「5.5」 and does NOT disambiguate it; on context (a 清华 researcher posting about ablation studies, tagged #科研日常) it reads as GPT-5.5, and the card labels that as an INFERENCE. What is certain is that it is not a Claude model — Anthropic ships nothing called 5.5. An earlier draft asserted 'GPT-5.5' flatly as though the note said it. ⚠️ NOTE ALSO WHAT CUTS AGAINST K3 in the same note, carried because a source used to rebut a vendor's weakness must be quoted whole: 「今天正式版干活 2h 用掉月度额度 10%」, 「和 k3 max(还只能 max)聊了一句消耗我月度 0.07%」, and 「看来是逼迫 199 用户升级到 699 吗。但是 699 性价比又很低」 — a THIRD independent witness that ¥699 is a live Kimi tier. SIGNED URLS EXPIRE: the durable handle is note id + author + body substring.
- 小红书 — xhs-manyouyou (search query, not a page) ⚠️ NO URL: the round-5 sweep did not recover a durable note id for this author, so this row deliberately carries NO link rather than a truncated profile path that resolves to nothing (an earlier draft carried `/user/profile/` with no id — a dead link). ⚠️ HONEST LIMIT, STATED: the round-5 sweep did NOT recover a durable note id for this author, so this row carries the AUTHOR and the QUOTE but NOT a note-level permalink — it is the weakest-anchored item in this ledger and is labelled so rather than given a URL it does not have. NOT COUNTED AS A SENTIMENT VOICE, in either §L (direction) or §N (astroturf) — which is also what removes the double-count risk a verifier flagged, since she appeared in both.
- B站 — www.bilibili.com/video/BV1UrKn64ECZ ⚠️ ADDED IN ROUND 5 — quoted in §Q4 and facts.md with no ledger row. @TwinkYou, 69 likes, on kate人不错's Java-backend test video: 「和预想的差不多,远强于4.8,稍逊于fable5」 ('about as expected: far stronger than 4.8, slightly behind fable5'). A lay Chinese comment landing unprompted on exactly the shape two independent eval houses measured — K3 BETWEEN the two Claudes. COMMENT, not spoken content: both B站 videos on this card return code 0 with an explicit empty subtitle list, so nothing spoken is quoted or characterised anywhere.
- 技术栈 / jishuzhan.net — jishuzhan.net/article/2078038773219860482 ⚠️⚠️ ADDED IN ROUND 5 AND SIMULTANEOUSLY DISQUALIFIED — CARRIES NOTHING ON THE CARD FACE. 孟健 (Meng Jian, ex-Tencent T11, ex-ByteDance tech lead), 「用 Kimi K3 交付一个真实项目后:很强,但还有不足」. Read in full via agent-reach 2026-07-20. For two rounds this source existed ONLY as a name inside another source's prose note — no URL, no quote, no date, no author, no platform, no ledger row — while being used as one of THREE 'independent first-hand contradictions' to Moonshot's own disclosed 'excessive proactiveness' weakness, i.e. rebutting a vendor's admission IN THE VENDOR'S FAVOUR. That is the highest-risk placement an unaudited source can occupy on this card. THE QUOTE IS REAL: 「方案摆完,它停下来等我拍板,没有替我做决定。」 ('with the options laid out, it stopped and waited for me to call it — it did not make the decision for me') is verbatim. BUT THE STAKE IS ALSO REAL, and it is in his own second paragraph: 「Kimi K3 正式发布了。这个结果,我终于可以公开说了:发布之前,我已经把 ShipSolo 微信小程序这个真实项目整个交给它」 — he had PRE-RELEASE ACCESS and was holding the writeup until launch day. That is a vendor-coordinated arrangement, disclosed by him and to his credit, and it is THE SAME STAKE CLASS as u/reaznval, the comped beta tester this card already excludes by name. Applying the card's own exclusion standard evenly means it comes out: the 'excessive proactiveness' contradiction count drops THREE -> TWO. Kept in the ledger with its URL and its stake so the exclusion is auditable — nothing is deleted for being inconvenient, and nothing is counted for being convenient.
- X (vendor account) — x.com/claudeai/status/2078302415804379218 ⚠️ VENDOR (Anthropic's own @claudeai account) — THE PRIMARY for the 2026-07-20 Fable 5 plan change, and primary for TIERS ONLY, never verdict-carrying. Text: 'Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.' Quoted here as relayed VERBATIM by simonwillison.net 2026-07-18 (read via agent-reach 2026-07-20) and independently by the-decoder, techtimes and thenewstack — i.e. four independent transcriptions agree. The same announcement ends the 50% boost to Claude Code weekly rate limits on the same date. FOUND IN ROUND 5 ONLY, after Verifier C found four card-face claims falsified by it — see facts.md §9-PRE.
- personal blog — simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/ Read via agent-reach 2026-07-20. Relays the @claudeai announcement verbatim and adds the line this card most needed: 'Important to note that users on the $20/month plan will still not have access to Fable 5 on that subscription. The Max plans are $100 and $200/month.' ⚠️ THE EMBARRASSING PART, RECORDED: the SAME AUTHOR is evidence[0] on this card for his 2026-07-16 K3 post. This post is two days later on the same blog and was never opened. He also asserts the reversal was forced by GPT-5.6 Sol and (lesser) Kimi K3 competition — a CAUSAL claim this card refuses to carry (facts.md already lists it as correlation-only).
- trade press — thenewstack.io/fable-5-permanent-subscription-access/ Read via agent-reach 2026-07-20. Carries an ANTHROPIC SPOKESPERSON ON THE RECORD: 'Fable 5 at 50% of usage limits is now a standard, permanent part of Max and Team Premium plans.' The strongest corroboration available short of the tweet itself, because it is a fresh attributed statement rather than a re-print. Used for the tier structure only.
- trade press — www.pcworld.com/article/3194944/fable-will-stay-in-claude-plans-but-not-for-everyone.html Read via agent-reach 2026-07-20. The cleanest price statement: 'Starting July 20 — today — Claude Max and Team Premium members will continue to have Fable 5 included in their subscription plans, at up to 50 percent of their weekly usage limits. Claude Max plans start at $100 a month (going up to $200/month), while Team Premium plans go for $100/month per seat.' And on the losing side: Pro ($20) and Team Standard keep access 'only via usage credits, which cost roughly as much as paying per token via the Anthropic API.' Author discloses he had not yet seen the $100 credit land on his own Pro account — a first-hand limit worth keeping.
- trade press — the-decoder.com/anthropic-slashes-claude-fable-5-limits-in-max-and-team-premium-and-pushes-pro-users-toward-api-pricing/ Read via agent-reach 2026-07-20. THE ONLY OUTLET THAT QUANTIFIES THE COMPOUNDING CATCH: 'The bonus usage phase ends the same day, cutting regular limits by 33 percent. Fable 5 will only be available at 50 percent of those already reduced limits.' That ~33% is this outlet's own arithmetic on the announcement, not a figure Anthropic published — carried on the card face as 'a baseline that shrank the same day', with the 33% attributed. Also asserts competitive causation (GPT-5.6 Sol, Kimi K3) which this card does NOT carry.
- trade press — www.techtimes.com/articles/320999/20260720/claude-fable-5-billing-splits-today-max-gets-it-free-pro-pays-per-token.htm Read via agent-reach 2026-07-20. Corroborates the $10/M input, $50/M output metered rate for Pro/Team Standard after the one-time $100 credit, and independently states the compounding effect ('Standard limits shrink first; then the 50% Fable cap applies against that reduced baseline'). Scored lower on reliability than PCWorld/thenewstack: the piece carries a fair amount of unattributed inference about Anthropic's compute economics. Used ONLY for the two figures, both of which are independently confirmed elsewhere.
- vendor product page — www.anthropic.com/claude/fable ⚠️ VENDOR — used for ONE fact: 'Claude Fable 5 is priced at $10 per million input tokens and $50 per million output tokens, with the existing 90% input token discount for prompt caching.' ⚠️ AND USED AS A NEGATIVE FINDING: read 2026-07-20, the page still says only 'available to Pro, Max, Team, and Enterprise users' and links a 'promotional access' support article — it does NOT state the 07-20 two-tier split. So the vendor's own marketing page was STALE relative to the vendor's own announcement, which is a reason the plan split is carried by the X announcement plus six independent outlets rather than by this page. Also the source for the mandatory 30-day data retention and the Opus 4.8 safety fallback.
- 小红书 — www.xiaohongshu.com/search_result/6a5b077800000000110116b7?xsec_token=ABDgkwAofr4XdFCpQXCtNMcw1G5Cz4EgHiuU5F6t_zp1k=&xsec_source= Title 「Fable 5 永不下线,大家却都在感谢 Codex」. Carries the strongest first-hand statement on this card that Fable 5 is genuinely ahead of K3 (「经过最近的高强度使用…Fable 5 还是显著优于 Kimi k3 和 GPT 5.6 Sol Ultra 的…就是有一股灵气」) — quoted on the card face at evidence_more[21], verbatim, and re-confirmed by Verifier C. Counted `neg` in the roster: the weighting axis is the model under review (K3), not the sentiment of the sentence. ⚠️ ROUND 5 — THIS NOTE WAS READ SELECTIVELY, AND IT COST THE CARD FOUR FALSE CLAIMS. Its FIRST paragraph is the 2026-07-20 Anthropic plan change: 「Claude 果不其然怂了,宣布从7月20日起,Fable 5 永久在 Max 和 Team Premium 计划中,但限额仍然是50%,Pro 和Team Standard 赠送100美元的调用积分作为补偿」. The builder took the quote from the bottom and never read the top. Ledgered here for the first time in round 5 — it was cited on the card face without a ledger row, which is itself a receipts gap. See facts.md §9-PRE.
- 小红书 — www.xiaohongshu.com/search_result/6a5d869d000000000100076f?xsec_token=ABg5BdPlstMxSPeFWIubk_ORtAYGfr1uUY_2qzuGbK7uc=&xsec_source= Title 「Kimi K3真实缺点汇总」. CARRIES ONE THING ON THIS CARD: a clear third-party statement of Moonshot's documented harness constraint — 「官方明确要求多轮对话和工具调用必须原样回传完整 assistant 消息,包括思考内容…并非所有现有 Agent 框架都能无缝替换」 (evidence_more[20], verbatim, re-confirmed by Verifier C). ⚠️ ROUND 5 — REMOVED FROM THE SENTIMENT COUNT, n 66 → 65. Re-read live: all six items are relayed third-party material (AA-Omniscience 33%→46% accuracy / 39%→51% hallucination; Artificial Analysis ~130M tokens vs a ~63M peer average; AA's $0.94/task vs GLM-5.2 $0.32 and DeepSeek V4 Pro $0.04; paraphrased official API constraints). NO hands-on use anywhere in it — the roster's own exclusion class, the one that already excludes Simon Willison. Kept as a source, dropped as a voice. Ledgered here for the first time in round 5. Reliability 0.6 because it is a relay: the underlying AA figures are not independently re-derived here and the card does not carry any of them.
-
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5e08dd000000000101e9dd?xsec_token=ABbmFhjLI2a_bslgp4785vw_dBuqI-lSouk6_454a3Q5Q=&xsec_source=
Jake — roster #37, weighted POS. Title 「也算是KIMI老用户了」, uid 5ee2cc460000000001000907, 2 likes / 1 collect / 0 comments. THE SINGLE FIRST-HAND Claude→Kimi CANCELLATION in the 33-note corpus, and the counter-example that falsified this card's own 「nobody is arriving from Claude」 (bug #42): 「说实话不是一直好用,几个月前的kimi code还是很渣的。但最近突然好使了起来,好使到我直接把claude退订了。」 and 「个人感觉,最近kimi cli的表现已经超过claude了。DS写代码还是差点意思。」 ⚠️ THE NOTE ENDS BY REVERSING THE CANCELLATION, and that is carried rather than tidied: 「退订了一个月,是被这家伙包庇以色列给气到了。气消了又订回来了」 (*"unsubscribed for a month, angered by this guy shielding Israel; the anger passed and I resubscribed"*). The referent of 这家伙 is genuinely ambiguous — it could be either vendor — and it is NOT resolved by guessing. Note the error direction: if it resolved against Jake the count would go to ZERO, which would STRENGTHEN this card's thesis, so the ambiguity is disclosed against our own interest. What it does not touch is the self-contained sentence that names Claude and states a cancellation. DURABLE KEY: note id 6a5e08dd000000000101e9dd + author Jake + the substring 「好使到我直接把claude退订了」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5be716000000000301fc9e?xsec_token=ABDgkwAofr4XdFCpQXCtNMcybUk1hMO1d4FwfWaDibFt0=&xsec_source=
CV — roster #38, weighted MIX. Title 「之前过的什么苦日子😭」, 5 likes / 2 collects / 2 comments. The WEAKER second Claude-departure, carried WITH the reason it is weaker: 「直接对Claude以及A➗祛魅了,乖乖改用温柔接住我的Codex,刚试了下Kimi3也是非常的不错啊 目前已经充了Gpt plus + Kimi Allegretto,价格一共才是Max的一半」. ⚠️ HIS TRIGGER WAS AN ACCOUNT BAN, NOT K3's MERITS — the note is tagged #claude封号 and opens 「给A➗送了几个月钱…往死里封」 — and HIS PRIMARY DESTINATION IS CODEX, with Kimi added alongside. Counting this as "arriving from Claude for K3" would be exactly the overreading this card exists to refuse, so it is counted mix and the card face says a second arrives "after an account ban with Codex as his destination". Reliability 0.6: a direction report, no measurement. DURABLE KEY: note id 6a5be716000000000301fc9e + author CV + the substring 「乖乖改用温柔接住我的Codex」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5d938d00000000010320c8?xsec_token=ABg5BdPlstMxSPeFWIubk_OQWByKt7VvF6xNBPfFKgdVs=&xsec_source=
0.999......=1 — roster #39, weighted MIX. Title 「深度用了几天K3和GPT-5.6,说几句大实话」, 4 likes / 1 collect / 9 comments. ONE OF THE THREE SUSTAINED-USE REPORTS whose existence falsified this card's 「no sustained-use report exists」 (bug #42), and it declares its own independence: 「kimi k3和gpt 5.6两个都深度使用,无广,纯实测」. CARRIES A NOVEL AXIS that appears on no leaderboard on this card — an instruction-following failure that is LANGUAGE-SPECIFIC: 「kimi K3 的硬伤:几乎全英文。思考链全英文,回答也大量夹英文。我明确提示"请用中文",没用,下一条照样英文思考。」 It corroborates the independent convergence on weak instruction-following from a direction none of the other voices took. Also 「一个共同感受:都变慢了。」 ⚠️ THE COMPARATOR IS GPT-5.6, NOT CLAUDE — his Program Bench figures (77.8 vs 77.6) are relayed, are not re-derived here, and the card carries none of them. ⚠️ Roster prints the handle as 0.999……=1 with CJK ellipsis; the account renders it with ASCII dots (0.999......=1). Same person, same uid 5ce7e55d00000000100017ff. DURABLE KEY: note id 6a5d938d00000000010320c8 + author 0.999......=1 + the substring 「无广,纯实测」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5cea39000000001c011880?xsec_token=ABXMBae_r4ctjdO-N9qWQ4ioTkTz29njzV_XEn23ddxlI=&xsec_source=
cuikq — roster #40, weighted MIX. Title 「真实工作环境kimi k3体验」, 15 likes / 4 collects / 6 comments. THE SECOND OF THE THREE SUSTAINED-USE REPORTS, and the one that is this card in miniature because BOTH HALVES COME FROM THE SAME SESSION. Economics: 「我是199rmb的订阅,日常工作强度,都是长任务,一下午两次hit 5h limit,大概只同时开了四个窗口」 and 「这个消耗速度个人体感快于gpt plus订阅下的5.6sol high」. Quality, in the opposite direction: 「任务的完成度,还有完成任务的思路,感觉基本可以标齐codex+gpt了,尤其是验证的思路非常明确扎实」. He also independently reports the English-CoT axis 0.999......=1 raises: 「只有大段的cot还是英文,比较难快速介入他的思考过程」. ⚠️ A BODY-FEEL COMPARISON (个人体感), not an instrumented one — carried as a report, not as a measurement, and the comparator is GPT, not Claude. Corroborates the ¥199 Kimi Code tier (§G) from a third first-hand user. DURABLE KEY: note id 6a5cea39000000001c011880 + author cuikq + the substring 「一下午两次hit 5h limit」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5dcfdb0000000004029f13?xsec_token=ABg5BdPlstMxSPeFWIubk_OdxKzHEQSLWNEOAfxA6srJE=&xsec_source=
Daveee — roster #42, weighted NEG. Title 「KIMI真实感受|评分80/100分|涨价后40/100」, 3 likes / 4 comments. A RANKED verdict from a user who states his own workload split (60% written/product work, 40% code) and names every tool he has used: 「1. 能力:Fable5>5.6sol>k3>opus4.8」 — which puts K3 ABOVE Opus 4.8 and BELOW both frontier rivals — and 「综合评价:优先考虑Codex/Claude…主要输在太慢和额度太少。」 ⚠️ WEIGHTED neg BUT IT IS NOT UNIFORMLY NEGATIVE, and the mixed part is carried: he rates K3 better than 5.6 Sol on written content and on front-end, and names two groups it genuinely suits (「需要开国内发票的/需要用合规AI的」 and 「就是想支持国产」) — the availability argument this card makes in switch_if. ⚠️ 个人体感 throughout: an ordered preference, no instrumented numbers, so it is a voice and not a measurement. Reliability 0.7 for the disclosed workload and tool set. DURABLE KEY: note id 6a5dcfdb0000000004029f13 + author Daveee + the substring 「优先考虑Codex/Claude」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5ce0d4000000000301ecdc?xsec_token=ABXMBae_r4ctjdO-N9qWQ4ip7Ho208wKb2Pf8ly6Hwbu4=&xsec_source=
A娃博物学家 — roster #43, weighted NEG. Title 「Kimi k3的抽卡式软件工程,不好接受」, 3 likes / 3 collects. THE SHARPEST RELIABILITY COMPLAINT IN THE CORPUS, and it is a three-way comparison he ran himself against Grok and Doubao-Seed-Evolving on near-equal workloads: 「K3在CLI里面依然磕磕绊绊,2次执行失败,我手动重启任务,最后耗时11个小时。」 and 「工程抽卡化了,我接受不了。」 ⚠️ NOTE THE STAKE DIRECTION, WHICH IS AGAINST THE FINDING: he opens by saying he SET OUT TO CLEAR K3'S NAME and delete his own critical post from the day before (「本意希望给KIMI正名的,好把昨天错误的帖子删掉」) and concludes 「我真难,无法推翻昨天的结论」 — a user trying and failing to falsify his own negative result. ⚠️ HIS EXPERT-ROUTING EXPLANATION (「892个专家的超稀疏选角」) is his own speculation and is NOT carried as a mechanism; note also that Moonshot's own launch post says 896 experts, not 892. DURABLE KEY: note id 6a5ce0d4000000000301ecdc + author A娃博物学家 + the substring 「工程抽卡化了」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a59f8d80000000001033227?xsec_token=AB0LmIDPbnCUqHZi9HPLl86qUCOdNvCtrd4GwUSzqQ0n0=&xsec_source=
学术猫猫(爱发论文版 — roster #44, weighted MIX. Title 「和Kimi K3吵架了12个小时之后,我总结出来了」, 5 likes / 1 collect / 20 comments. A NOVEL AXIS — who the model suits BY DISCIPLINE, from ~12 hours of non-coding use: 「如果是人文社科的朋友,别碰,会变得不幸。如果是理工类的,勇敢冲。」 and 「有点像个智商为1000情商为0的呆子了」. Her mechanism is stated and is checkable against the card's own framing: a coding benchmark optimises for deterministic convergence, so the edge-case pedantry that wins there transfers badly to statistical inference (「在代码里,一个未被处理的 edge case 能让整个系统崩溃;在社科里,揪着 2σ 之外的尾部噪声不放,只会导致分析瘫痪」). ⚠️ CARRIED WITH ITS LIMITS: this is NOT a coding report and NOT a Claude comparison — no rival is named — and the axis (情商 / humanities fit) is not measured by any benchmark on this card. Reliability 0.6 for that reason, neutrality 0.85 (no stake found; she also records K3 conceding the point to her, which cuts against her own thesis). ⚠️ THE ROSTER PRINTS THE HANDLE TRUNCATED as 学术猫猫; the account reads 学术猫猫(爱发论文版, uid 5e6876d700000000010088b5. DURABLE KEY: note id 6a59f8d80000000001033227 + author 学术猫猫(爱发论文版 + the substring 「有点像个智商为1000情商为0的呆子」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a59e3e4000000000402a5b2?xsec_token=AB0LmIDPbnCUqHZi9HPLl86qKXiVT7TE5eH6KYOH7kI00=&xsec_source=
孤岛 — roster #45, weighted POS. Title 「入坑Kimi K3,今天给它派了个坑活。」, 8 likes / 2 comments. THE STRONGEST POSITIVE AGENTIC REPORT IN THE 小红书 CORPUS, and it is carried at full weight precisely because most of this corpus runs the other way: a real migration task (7 Obsidian docs → FlowUs) in which K3 span up four subagents, hit an API-induced fault that duplicated 9,000+ table rows, and RECOVERED WITHOUT HIM — 「中途 API 网络抖动,表格行重复刷了 9000 多条……我正想上手收拾,它自己查到原因(重试非幂等),自己写脚本把重复行清干净重新传了。」 His summary: 「我全程就干了三件事:提需求、确认方案、看戏。」 ⚠️ LIMITS, STATED: a single self-reported run, no rival ran the same task, no timings, and the fault it recovered from was one it had itself induced. It is a voice and an existence proof, not a measurement — and it is a first-hand contradiction of the reliability complaint at xhs-awa-bowuxuejia, which is why both are carried. DURABLE KEY: note id 6a59e3e4000000000402a5b2 + author 孤岛 + the substring 「我全程就干了三件事:提需求、确认方案、看戏」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5cea0b000000000c033c02?xsec_token=ABXMBae_r4ctjdO-N9qWQ4iuOjgfUblXUgWRD18Xphk24=&xsec_source=
阿西吧 — roster #46, weighted NEG. Title 「被Kimi3杀猪了」, 5 likes / 1 collect / 10 comments. ONE OF FIVE FIRST-HAND QUOTA-EXHAUSTION CORROBORATIONS from this sweep, and the one on the LARGEST plan: 「今天醒来验收,一个子代理卡死,然后任务没结果,账单有结果,把我这个月的额度直接耗光了!我是20倍的额度计划啊!」 — a stuck subagent burning a whole MONTH of the 20× tier with no deliverable. Also a duration comparison: 「一个小任务,同等任务codex一般5-10分钟搞定的,它折腾了一晚上」. ⚠️ THE COMPARATOR IS CODEX, NOT CLAUDE, and 「同等任务」 is his own judgement of equivalence, not a controlled one — carried as a report, not as a measurement. Neutrality 0.8, reliability 0.65: an angry post with no screenshot (contrast xhs-blueraincoat, which scores higher because it has one). DURABLE KEY: note id 6a5cea0b000000000c033c02 + author 阿西吧 + the substring 「我是20倍的额度计划啊」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5a305e0000000010029f30?xsec_token=ABOF6h4dhCuHuZA1ibrm5Jk4gBU-OBod0MN9-S6YiF33g=&xsec_source=
小红薯63975696 — roster #47, weighted NEG. Title 「kimi k3 消耗巨快」, 4 likes / 1 comment. THE MOST INSTRUMENTED of the five quota corroborations, because he names his meter: 「我200套餐,用了不到一个小时,还不是连续用,5小时限额打满,周限扣了25%多,ccswitch 统计不到500w token,非常离谱。」 — under an hour, non-continuous, 5-hour cap exhausted and >25% of the WEEKLY cap gone, against under 5M tokens counted by the ccswitch monitor. That third-party token count is what makes it a measurement rather than an impression, and it independently reproduces the mechanism at xhs-zhongmeng (agent loops are cache-read dominated, so real cost detaches from headline token volume). ⚠️ 「200套餐」 is not disambiguated (¥200-class Kimi plan) and NO CLAUDE COMPARISON is made anywhere in the note. Default handle (小红薯 + digits) = an account that has never set a nickname; it does not imply a throwaway, but nothing about the author is checkable either, hence reliability 0.65. DURABLE KEY: note id 6a5a305e0000000010029f30 + author 小红薯63975696 + the substring 「5小时限额打满」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5c51f0000000002201a2e9?xsec_token=ABXMBae_r4ctjdO-N9qWQ4iqyOm6A2CqNa3r2lVqi6QIA=&xsec_source=
落梦 — roster #48, weighted NEG. Title 「Kimi干到一半提桶跑路,好绝望🥲」, 6 likes / 19 comments. The third of the five quota corroborations: a ¥50-tier user whose allowance ran out MID-TASK, during troubleshooting of the very thing K3 had just told him to do — 「排查问题的过程中,它给我来了一句额度用完了」 — on 「k3极致模式」 (the max reasoning-effort setting), which he says he had barely used before. ⚠️ THE WEAKEST-ANCHORED OF THE FIVE and scored accordingly (reliability 0.6): the smallest plan, no token figure, no timings, and his wider complaint is about K3's PREDECESSOR (kimi 2.6) and about server load, not about the comparison this card makes. ⚠️ HIS COMPARATOR IS GPT-5.6 SOL, NOT CLAUDE. It is carried because five independent users hitting the wall on four different tiers is the finding, not any one of them. This row supplies the durable key the round-5 receipts left blank for him. DURABLE KEY: note id 6a5c51f0000000002201a2e9 + author 落梦 + the substring 「它给我来了一句额度用完了」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5cea5b000000002201a2f8?xsec_token=ABXMBae_r4ctjdO-N9qWQ4ijFnSt_VhGAKtm7dweZDK4M=&xsec_source=
Unibox0123 — roster #49, weighted NEG. Title 「压线买了 Kimi 订阅…」, 11 likes / 3 collects / 9 comments. THE BEST-ARITHMETIC quota corroboration: he converts his own burn rate into the number of sessions the plan actually buys — 「几个小时消耗掉总用量4.2%,算下来满负荷用24轮(5h/轮)。但是按周来算,一周只能用5轮。」 (~24 five-hour rounds in a month, but only 5 a week once the weekly cap binds). That is the same shape as the card's subscription-economics finding, derived independently on the Allegretto tier. He also independently records the purchase freeze (「下午买了…没想到晚上官方就说不再开放购买了」). ⚠️ AND HIS QUALITY VERDICT IS POSITIVE — 「体感上K3发挥比较稳定」 — carried because a source used against a product must be quoted whole. ⚠️ NO CLAUDE COMPARISON; the percentages are his own console readings, unverifiable by us. Highest reliability of the five (0.75) because it is the only one that shows its arithmetic. DURABLE KEY: note id 6a5cea5b000000002201a2f8 + author Unibox0123 + the substring 「几个小时消耗掉总用量4.2%」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39). -
xiaohongshu (CN) — www.xiaohongshu.com/search_result/6a5b40b5000000001102c577?xsec_token=ABDgkwAofr4XdFCpQXCtNMcyBlMxsSc2A411bm-b7z2ng=&xsec_source=
Xaiver — roster #50, weighted NEG. Title 「Kimi K3性价比实测结论…」, 27 likes / 14 collects / 59 comments — THE HIGHEST-ENGAGEMENT of the five quota corroborations. A self-described heavy AI-coding user who subscribed to a rival Pro plan on the 14th and to Kimi on the 17th, comparing the two on the same codebase: 「在我的大型 code base 里面,一轮任务甚至没跑完就中道崩殂了」 and 「在性价比方面,两者的体感真的差异非常非常明显」. ⚠️ THE DECISIVE LIMIT, AND IT IS WHY THIS SCORES 0.65 AND NOT HIGHER: HE NEVER NAMES THE RIVAL. 「某厂pro」 is literally "a certain vendor's Pro" — it may or may not be Claude, and this card does NOT treat it as a Claude comparison. What it carries is the unambiguous half: a large-codebase task that died mid-run on quota. The engagement (59 comments) is NOT used as a signal — see xhs-kimi-official for why engagement carries nothing here. DURABLE KEY: note id 6a5b40b5000000001102c577 + author Xaiver + the substring 「一轮任务甚至没跑完就中道崩殂了」. Read 2026-07-20 via agent-reach; likes/comments are the counts read that day and 小红书 counts drift. ⚠️ SIGNED URLS EXPIRE. The durable handle is note id + author + body substring: search 小红书 for the author or the substring and open the SIGNED (?xsec_token=…) url the search returns — a bare /explore/
or unsigned /search_result/ is rejected (bug #39).
Stuck on a different switch?
If the card doesn't exist yet, request it — free, like everything here. Full sweep, weighted verdict, and one email the moment it's published.
Request a card