Switch card · AI CODING MODELS
Kimi K3 Claude
CONFIDENCE: LOW· updated 2026-07-20· n=65 independent, 28 sources· vendor content discounted

On the only two independent same-harness boards, K3, Opus 4.8 and Fable 5 are inside each other's noise. Terminal-Bench 2.1, one evaluator, one scaffold: K3 85.0% · Fable 5 84.6% · Opus 4.8 84.6% — K3 ahead of both by one task instance. DeepSWE, one scaffold: Fable 5 70%±4 · K3 69%±5 — overlapping intervals, at $21.63 vs $4.65 per task. But that cost gap is an API-list-price gap. The practitioners who ran both on subscriptions report it vanishing — “K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan” — and it is now measured — by two people, and they are not the same person: u/djdante reports that at $20 a side, K3 consumes 4× the weekly quota of Fable 5 High for the same task (⚠ his own words: he ran those tests “for my YouTube channel”, and no artefact, harness or task set is published — disclosed because it costs us, not because it helps) 🕐 and read the date: that was measured 2026-07-15, against a $20 Claude Pro plan that still included Fable 5. Pro stopped including it on 2026-07-20. The measurement stands as made; the configuration does not. To run it as of 2026-07-20 you would need a $100 Max plan on the Claude side; u/ahnyudingslover, replying directly to him on a production codebase, adds that output quality “was close, which was shocking” (a direct reply opening “I can attest to this as well” — corroboration in-thread, not a second independent measurement, and also a $20-tier report carrying the same pre-07-20 dating caveat). So if you pay by plan rather than by token, the cheap-K3 case is not merely unproven — it points the other way. (It reverses again at $100, where the same complaint becomes “very generous”. Tier matters.) So the thing that actually separates these models is not the score. It is what you can buy: K3's weights aren't out, its licence is unannounced, and Moonshot paused new subscriptions on 07-19 — while on the Claude side, from 2026-07-20, Fable 5 is a permanent part of Max and Team Premium only, and there at 50% of each plan's usage limits. Max is an individual plan — it starts at $100/month and runs to $200. Claude Pro ($20) and Team Standard lost their included Fable 5 on that date: they keep it only through usage credits billed at API rates ($10/M in, $50/M out) after a one-time $100 credit. And the cap compounds — the 50% boost to Claude Code weekly limits ended the same day, so the half-slice is measured against a baseline the-decoder puts ~33% lower. “Just use Fable 5” is still not a $20-plan answer — it is a $100-plan answer, at half of a limit that just shrank. Both sides have an access problem, and neither has a settled quality gap. That is a timing answer. Run K3 as a second seat on the shapes it demonstrably wins. Leave the loop you ship from where it is. ⚠ And read our confidence precisely, because the two halves differ: this card is LOW because the quality and harness questions are wide open — the K3–Fable gap is one instance in 267 on one board and inside the error bars on the other, pointing opposite ways, and no independent test in any language — across every source in this sweep — has run the same tasks in both a Kimi and a Claude-shaped harness. But the subscription-economics finding specifically is close to settled: one quantified $20-vs-$20 ratio plus first-hand corroboration on six platforms in three languages. (That used to read “two quantified measurements”. It was an over-claim: djdante gives a ratio, ahnyudingslover gives only Fable's own consumption and no K3 figure. One ratio, corroborated in-thread.) LOW is a judgement about the verdict, not about that finding. ⚠️ And a warning about that word “settled”: a third verifier found this card falsified on its own publication date — on the access half, the half we were calling near-settled — because Anthropic changed which plans include Fable 5 on 2026-07-20 and we had not read it, in two sources this card already cites. It is corrected above. Treat “near-settled” as a statement about a moving object, and check the plan tables yourself before you spend anything.

The fact that decides it: the claim doing the most work in launch coverage — that K3 loses to Claude Fable 5 — comes from Moonshot's own blog, and the independent same-harness data does not reproduce it. Artificial Analysis ran K3, Fable 5 and Opus 4.8 through the same Terminus 2 harness: 85.0 / 84.6 / 84.6. Datacurve ran all three through the same mini-swe-agent: 69±5 / 70±4 / 59±2. A vendor conceding a gap that neutral measurement cannot find is a marketing fact, not a finding — and this card refuses to let it carry the verdict. What survives contact with independent data is cost and availability, not capability.

The decision

CHOOSE CLAUDE IF →

  • Your work is greenfield frontend, visual or 3D — Frontend Code Arena #1 at 1679, “surpassing Claude Fable 5” and #1 in 6 of 7 domains (Arena's own post); the relayed leaderboard puts Fable 5 at 1631 and Opus 4.8 (Thinking) at 1562. Corroborated by testers otherwise hostile to the model
  • Cost per finished task binds, you bill at API rates, and you can get a seat — same-harness DeepSWE: $4.65 vs Opus 4.8's $13.22 and Fable 5's $21.63, on fewer output tokens than either (81k vs 135k vs 119k). This advantage is an API-list-price advantage and does not survive onto a subscription — see the stay-put column
  • You want frontier-ish coding on a cheap plan — because from 2026-07-20 the cheapest Claude plan that includes Fable 5 is Max, at $100/month (Max 5× $100 · Max 20× $200 · Team Premium $100/seat), and even there it is capped at 50% of that plan's usage limits, against a weekly baseline that shrank the same day. On Pro ($20) and Team Standard, Fable 5 is now credits-only at API rates — $10/M in, $50/M out — after a one-time $100 credit. Opus 4.8 needs Pro or above. K3 at $3 / $0.30 cached / $15 per MTok is a third of Fable 5's $10 / $1 / $50
  • You'll run it in Kimi Code, or a harness you have verified echoes thinking history back — the vendor's own words are that quality “may become highly unstable” otherwise
  • You need security analysis a Claude model refuses — a private, un-gameable cyber eval rated K3 “best price/recall” (best value for recall, not best recall) and recommended it for continuous analysis, with Fable 5 at a 100% refusal rate

CHOOSE KIMI K3 IF →

  • You want the gap to be real before you move — K3's edge over Fable 5 is one task instance on one board and a reversal inside overlapping error bars on the other. Nothing here clears noise in either direction
  • Your work is maintenance on an existing large repo — an independent repair harness put K3 last of 7 (79%), with both Claude models ahead of it
  • Your harness is Claude-shaped and you can't verify the thinking-history echo — one live session-wedging bug that names Kimi — but the PR names Moonshot/Kimi and ZAI and carries both provider labels, so it is a two-provider wire-format bug, not K3-specific. Anthropic-style thinking blocks are one of the entry paths it enumerates, not the cause. Hit four times in one day. (This line used to read “three live bugs in three days”. Checked one at a time against the GitHub API on 2026-07-20: the second was closed by its own filer nine minutes after he opened it, and the third is a gateway bug hitting fifteen models with K3 as one list entry. One bug, not three.) The sharpest first-hand report is not on GitHub at all — a B站 user wiring Kimi into Claude Code via ccswitch with the model set to k3, asking why k2.6 is what actually gets called. No answer to him appears in the 107 replies we read (pages 1–3 of 372, hot-sorted) — the video's full comment set is login-walled, so we cannot say nobody answered, only that no answer is in what we read
  • You're on a subscription, not list API — this is the single strongest argument against the whole cheap-K3 case, and it is first-hand: “Kimi K3 burns my Kimi subscription as quickly as Fable burns my Anthropic plan — i observe no efficiency benefit” (@kunchenguid, on the $40 plan). Corroborated independently: a measured ¥199 Kimi Code plan came out at roughly a Claude Pro Fable allowance (🕐 measured 2026-07-17 — that comparator is gone: Claude Pro stopped including Fable 5 on 2026-07-20), and a 94% cache-hit user still burned the 5-hour window in two tasks. Both vendors meter on a rolling 5-hour window plus weekly caps
  • You need the weights, a licence, or a seat — as of 2026-07-20 none of the three existed: no HF repo, licence unannounced, new subscriptions paused

Works with your setup?

vs Opus 4.8
same harness
vs Fable 5
same harness
Claude-shaped
harness
Best model on
a cheap plan
Weights /
licence
Can you buy
a seat?
Kimi K3⚠ +1 instance / +10 pts⚠ tie, ~1/4 cost⚠ run, never measured⚠ ¥199 for 1M ctx✗ neither yet✗ paused 07-19
Claude (Opus 4.8 / Fable 5)peer, 2.8× cost⚠ tie, ~4× cost✓ it's native✗ Fable needs $100+ Max, at 50% limits✗ closed, stated✓ on sale

Sentiment — independent voices only

REBUILT FROM AN AUDITABLE PER-VOICE ROSTER ON 2026-07-20, and the number went DOWN TWICE: 75 → 66 → 65. Two independent verifiers reported that the previous n=75 could not be recounted — there was no per-voice ledger anywhere in the receipts, so the figure could not be re-derived and double-counting could not be excluded. One reconstructed it by hand and could only reach “≈73–77”. A count nobody can re-derive is not a real count; it is an assertion — and iron rule #3 says real counts. So there is now a 65-row roster, one line per voice (handle · platform · source id · weight) in the receipts, with the counting rule stated so someone else can apply it and get the same answer. Nine entries came out — analysts, benchmark relays and astroturf-allegation voices that had been swept into a sentiment count they never belonged in. Then a third verifier found a tenth: 小红書 老K, whose note is titled 「Kimi K3真实缺点汇总」 — a compilation relaying Artificial Analysis figures with no hands-on use anywhere in it, which is the roster's own stated exclusion, the same one that (correctly) excludes Simon Willison. The rule had not been applied evenly, and the miss ran in the direction that kept n higher. 66 → 65 (neg 34 → 33). He is still on this card and in the ledger, carrying what he can carry: a third party stating the vendor's documented harness constraint. No evidence was deleted: every one of those sources is still on this card. They are simply no longer counted as voices, because they are not. The card now under-claims its corpus rather than over-claiming it — the only direction we are permitted to err in. How to read the balance: negative skews high because the loudest cohort in the sweep window (2026-07-16 launch → 2026-07-20 sweep) is subscribers hitting quota walls, not people judging output quality — read it as a signal about economics and availability, not capability. Counter-weight: named 小红书 users allege coordinated paid posting (「铺天盖地的商单」), and the single highest-engagement K3 post on that platform is Moonshot's own account — so we weight GitHub bug reports and same-harness numbers above sentiment volume, and use engagement as a signal nowhere. One item — marked ⚠ live page returns HTTP 567 — cannot be loaded, though its text is independently corroborated by a third-party index.

Weighted evidence

Show all 28 sourcesShow fewer

Why confidence is LOW

three same-harness rows: K3, Opus 4.8, Fable 5 Chinese + Japanese sweep carried the cost finding $20 vs $20 on 07-15: K3 burned 4× the weekly quota — Pro dropped Fable 07-20 Chinese corpus: the exit from Claude runs to Codex, not K3 4 DAYS OF PUBLIC USE AT SWEEP (07-16 → 07-20) — only 3 sustained-use reports in 33 notes weights unreleased, licence unannounced new subscriptions paused 07-19 every gap vs Fable 5 is inside the error bars page one of Google: zero independent reviews

Coverage — what we read, what we skipped

⚠️ THIS FOOTER WAS REWRITTEN ON 2026-07-20 BECAUSE MOST OF IT WAS WRONG — and wrong in our own favour's opposite direction. Rounds 1–3 of this card were swept with plain web fetching. The research tool our own sweep contract mandates (agent-reach) was never called. The result was not a card that over-claimed. It was a card that invented weaknesses it does not have: four 小红书 sources were published as “login-walled, not re-verifiable” when they read perfectly well (小红书 needs a signed URL; we kept fetching the unsigned one, got a JavaScript shell, and called it a login wall); B站 was characterised without ever being read; 掘金 was written off on a search-syntax artefact; and 知乎 was described as empty when in fact its search is CAPTCHA-walled. A card that invents a weakness is as dishonest as one that invents a strength — it just feels like modesty at the time. Everything below is what the re-sweep actually found.
READ: Eval houses — Artificial Analysis (Terminal-Bench 2.1 same-harness, model comparison page, GDPval-AA v2), DeepSWE/Datacurve, tbench.ai (as a guard, not a number) · GitHub primary bug reports via the API (hermes-agent #67386, opencrabs #616, opencode #37635 + a five-issue prior-generation cluster, kimi-cli #1994) · Hugging Face API + K2.7 licence text · Vendor pricing pages, both sides (platform.claude.com pricing docs, claude.com/pricing, Anthropic Pro/Max support articles, Moonshot's blog + three doc sites) — ⚠️ read as primary for prices and tiers only, scored neutrality 0.0–0.2 in the ledger, never allowed to carry a verdict · Reddit (r/ClaudeCode, r/kimi ×2 incl. the 60-comment subscription-comparison thread, r/codex, r/OpenAI, r/hermesagent) · Hacker News — three threads, cited separately with real counts: 48935342 1,206 comments / 2,090 pts, 48960218 598 / 623, 48947717 221 / 403 · X practitioners · simonwillison.net ×2 posts (the K3 post AND the 07-18 Fable-permanent post — the second added only in round 5, see below) · Trade press on the 2026-07-20 Anthropic plan change (thenewstack, PCWorld, the-decoder, techtimes, androidauthority, aipricing.guru) plus @claudeai's own announcement · 中文: V2EX ×4 threads, 小红书 ×33 notes (this READ list said ×9 until 2026-07-20 — the pre-round-5 figure, never updated when the saturation sweep tripled it, while three sentences further down this same footer called it a 33-note sweep. 9 is 33 minus the 24 that were new. Corrected under this card's own published rule: an inaccurate method statement is inaccurate in either direction.), 阅微堂 (zhiqiang.org), 人人都是产品经理, 知乎, 掘金, B站 ×5 uploaders read of 13 found · 日本語: note.com, Qiita, PC Watch · YouTube + transcripts · both Google SERPs.
WHAT THE LEDGER BELOW CONTAINS, AND WHAT IT DELIBERATELY DOES NOT. Every one of the 65 voices behind n=65 now has a ledger row, which is the point of the ledger and was not true until 2026-07-20: 14 of the 65 were uncountable from this page — 13 小红书 notes from the saturation sweep with no row at all, plus one roster line citing a price instead of its note id — and the roster that lets you recount n is in the receipts, which this site does not publish. So the ledger was the only audit surface a reader had, and for 14 voices it was empty. Each of the 13 was re-found through agent-reach (search → signed URL), re-read end to end, and ledgered with its note id, author, a distinctive body substring and its weight; a further 5 roster lines pointed at ledger ids that did not exist and now cite the real ones. Five items this sweep read are deliberately NOT in the ledger, named here so the omission is auditable rather than silent: two 小红书 notes that are image-only — the body is a picture, so nothing is quotable from them (「Kimi K3 风评大变」 6a5c99bd…, 102 likes; 矩元 6a59ee33…, 9 likes) — and three notes dated 2026-07-21 (6a5e4cea…, 6a5e5318…, momo 6a5e50c6…), quarantined because a card dated 07-20 may not cite 07-21, which is a rule about principle and not about their merit. They are listed so the next sweep collects them instead of rediscovering them.
CORRECTED ON THE RE-SWEEP (2026-07-20) — every one of these was a claim this card made, and got wrong: (1) The four 小红书 “not re-verifiable” flags are GONE. All four notes were read via agent-reach 2026-07-20 (search → signed URL), and each was read end to end — these are single-page notes with no pagination and no collapsed replies, which is why an exhaustiveness claim is safe here and was not safe on the Reddit thread. To re-check one yourself: search the author or a phrase, then open the signed link (?xsec_token=…) the search returns — a bare /explore/<id> is rejected. Signed links expire, so for each note the receipts record the durable key instead: note id + author + a distinctive body substring. (2) A quote of ours was not verbatim. 终梦's conclusion was printed inside 「」 under a heading reading Verbatim as 「日常杂活 DeepSeek,最难的活给 K3」. He wrote 「我的方案:日常杂活 DeepSeek 焊死,最难的活给 K3 单独开小灶。」 A tidied-up abbreviation presented as a quotation is a fabricated quote that happens to be true; it is fixed everywhere and logged with the same weight as this card's round-1 fabrication. (3) “Three live session-wedging bugs in three days” was one bug. opencrabs #616 is closed — by the repo owner who filed it, nine minutes later; opencode #37635 is an opencode-go gateway bug hitting fifteen models with kimi-k3 as one list entry. Only hermes-agent #67386 is live — and on re-verification it is not K3-specific either: its body opens “Strict OpenAI-compatible providers (Moonshot/Kimi, ZAI)…” and it carries provider/kimi and provider/zai, so it is a two-provider wire-format class bug, not a K3 defect. Zero of the three original bugs is both live and K3-specific. On its labels: 6 of the 9 are ordinary repo taxonomy (type/bug, comp/agent, provider/kimi, provider/zai, P2, area/sessions) and only 3 carry the sweeper: prefix, so “maintainer-labelled” is now “repo-labelled”. kimi-cli #1994 was dated 2026-06; it was filed 2026-04-22. (4) There is no “~1,850-comment HN thread”. That number was two separate threads added together (1,206 + 598) and then described as one. Three real threads, listed above. (5) “Every V2EX reply re-read line-by-line” was not true. V2EX's API has an SSL failure, so those threads come through Jina Reader — whose render is lossy and silently drops replies (eleven missing from t/1228031 alone). Every V2EX quotation on this card did check out; the claim of exhaustive reading did not. (6) Two comment counts were stale (幸运的蜗牛 59→75, blueraincoat 36→49) and one author was misattributed: “V2EX @blueraincoat” — there is no V2EX evidence for him at all; he is 小红书 only. (7) 阅微堂 was fine all along — it is simply the blog name for zhiqiang.org, which has been in the ledger with a live URL throughout. It read as unfindable only because this footer printed the Chinese name with no domain beside it. Named properly now, not deleted.
UPGRADED, NOT DOWNGRADED: x.com is no longer login-walled to us. All four x.com items were re-verified verbatim through agent-reach, independently of the operator's browser capture — so they are now re-verifiable by you, not merely “captured from a session we hold”. The re-check also confirmed the card's own caution: the @arena post contains 1679 and “surpassing Claude Fable 5” and contains neither 1631 nor 1562, exactly as the card says. · 51CTO now has two witnesses. Its page still returns HTTP 567 (a bot-wall body; agent-reach gets an empty body) — but a third-party semantic index carries the full article text and matches both quoted strings verbatim, by a route with no access to anything of ours. It is the only genuinely hard-to-check source left on this card, and still its weakest item.
GAPS, STATED HONESTLY AND QUOTED WHERE WE HAVE THE ERROR: 知乎 enumeration was never achieved. Article pages read fine; zhihu.com/search returns 「系统监测到您的网络环境存在异常」 and “This page maybe requiring CAPTCHA”. So we can say what the pages we reached contain — the one substantive item was Moonshot's own launch blog, reposted, which carries nothing — but not what 知乎 contains. That is narrower and truer than the “returned nothing” this card used to print. · 微信公众号 via Sogou remains inaccessible — a real gap, disclosed. · B站 spoken content is not cited anywhere on this card. Both videos have no subtitle track — the player API answers code 0 (OK) with an explicit empty subtitle list, which is a definitive negative, not a fetch failure. So every B站 item here is existence, uploader-written description, or comment, never speech. And the largest video's full comment set is login-walled (「登录后查看 372 条评论」): pages 1–3 of the public, hot-sorted reply API were read — 107 of 372 replies, under a third, not all of it, and hot-sorting favours high engagement. (This footer previously said “the twenty comments read”, which under-reported the sweep and implied a ceiling never tested. An inaccurate method statement is inaccurate in either direction.) · Zero self-hosters of K3 in any language. · Zenn, はてなブックマーク and Korean (Velog/tistory) had no hands-on K3 coding writeups. · The “excessive proactiveness” check still finds nothingfive phrasings through an authenticated Reddit session returned no on-topic result, and the three HN threads grepped for proactive/improvis/overeager return zero. (Stated that way deliberately: with semantic search a non-match comes back full of unrelated posts, so the old “empty result set” overstated how clean the null was.) Meanwhile it has two first-hand contradictions that survive a stake check, and zero corroborations — one of them a user documenting K3 retracting its own conclusion unprompted. ⚠️ It was three until 2026-07-20, and we removed one ourselves: the ShipSolo production writeup is real and readable, but its author's own second paragraph discloses pre-release access held under embargo (「发布之前,我已经把 ShipSolo 微信小程序这个真实项目整个交给它」). That is the same stake class as the comped beta tester we already exclude by name, and it was being used to rebut a vendor's own disclosed weakness in that vendor's favour — the highest-risk seat an unaudited source can occupy here. It is kept in the ledger with its URL and its stake, and carries nothing on this card. Neither surviving contradiction compares against Claude: Z3PH1NUE's note writes only 「5.5」 and does not disambiguate it — on context that reads as GPT-5.5, which we label an inference; what is certain is that it is not a Claude model.
REFUSED, AND WHY (iron rule #1 — readers pay, vendors never do): AICodeKing (YouTube) discloses “Sponsored by Moonshot AI” — one of the highest-reach English K3 “tests”, and it carries nothing here. Dubibubi carries an affiliate offer for the product under review. 小红书 Kimi智能助手 is Moonshot's own account, and its launch post (8,022 likes read 2026-07-20; it was 8,016 when first captured, and the figure is a timestamp rather than a constant) is the highest-engagement K3 item in the whole 小红书 corpus — which is the concrete reason engagement volume is used as a signal nowhere on this card. Three Reddit commenters were excluded from the count for stake: a comped Kimi beta tester, someone who posts for Cline, and one whose comparison table his own comment says came from asking GPT. The page-one SEO/reseller cluster (kie.ai, myclaw.ai, layer3labs.io, llm-stats.com, codersera, vallettasoftware, digitalapplied) was read for cross-checkable facts and discounted for the verdict: 4 of 10 page-one results for “Kimi K3 vs Claude” and 6 of 10 for “vs Claude Fable 5” are commercial or affiliate properties, and they uniformly republish Moonshot's KimiCode-harness 88.3 as a like-for-like win — the most propagated error about this model. Named 小红书 users allege the pro-K3 wave was bought: 「Kimi K3凌晨上线,随后就是铺天盖地的商单,都快把真实用户的声音淹没完了」, seconded independently by another. We cannot verify that; we can and do weight bug reports and same-harness numbers above sentiment volume because of it.
WATCH OUT: The number this card most wants still does not exist — but the silence is smaller than we said. No independent test has run the same K3 tasks in both Kimi Code and a third-party harness, so the harness penalty itself is still unquantified — though K3 has now been run publicly inside Claude Code (AI超元域, B站, 2026-07-17) and inside a shared agent harness alongside rivals (Token就是词元, B站, 2026-07-18). Neither closes it: the Claude Code run used different tasks on the Kimi side, the shared-harness run isolates the model, not the harness, and the one YouTube head-to-head moved two variables at once (model and harness, plus an undisclosed Claude tier against a $19 plan). · The ¥699 Kimi-client price is contested — one first-hand user says ¥699 for 1M context, another the same week says ¥199, and nothing adjudicates. Only the Kimi Code tiering is corroborated by both (¥199 = 1M · ¥99 = 256k · ¥49 = no K3). · The “max” in the benchmark rows is a reasoning-effort setting, not the Claude Max plan — and the Fable row on Terminal-Bench is additionally an Opus 4.8-fallback configuration. Both same-harness gaps against Fable 5 are smaller than their own error bars and point in opposite directions; treat “K3 beats Fable” and “Fable beats K3” as equally unsupported. · Sustained maintenance on a large existing repo is unmeasured in every language. · And read the direction before you read the verdict — as a COUNT, because the absolute we used to print here was FALSE. This footer said “nobody is arriving from Claude”. One person is (小红书 Jake, 「好使到我直接把claude退订了」, cited above), and a second arrives after an account ban with Codex as his destination. What a 33-note sweep actually supports: the dominant exit vector from Claude in the Chinese corpus is Claude → Codex, with Kimi a secondary or supplementary seat, and exactly one first-hand Claude→Kimi cancellation surfaced across those 33 notes. Most arrivals come from Codex, GLM and DeepSeek; one recommends 「100刀gpt pro, 20刀claude,gpt当主力」; one says Claude and GPT resetting allowances makes him 「真的更想直接去用它们」. Claude is the seat these practitioners are trying to keep, not the one they are leaving — which is what the ⇄ is for.
⚠️ WHAT ROUND 5 REMOVED FROM THIS CARD, after TWO independent adversarial verifiers both refused to pass it. Both agreed the verdict, the arithmetic and the vendor hygiene were sound. Both refused the same thing: one recurring defect class — asserting a characterisation of material that was not read, or cannot be reached. That is bug #39's exact shape, and it was committed several more times inside the round convened to close it. (1) “Read in full” on the Reddit thread was FALSE — that render collapses 12 reply subtrees; the claim is replaced everywhere by a method and its limits. (2) A B站 quote could not be located in the description, in 184 comments including all eight uploader replies, or in the reply API — it is deleted and treated as fabricated, along with the claim resting on it. (3) Two universal negatives on this face were falsified by a 33-note saturation sweep — “nobody is arriving from Claude” and “no sustained-use report exists”. Both are now counts. (4) Our own cuts register was certifying a TRUE sentence as fabricated (「官方建议64卡起步,普通玩家基本不用想」 is 李卓's, on 51CTO, a source we already cite). A false accusation of fabrication, living inside the artifact whose whole job is to prevent fabrication, is worse than the original error — which was a misattribution. (5) Three 小红书 links on this face did not open, in the exact form this footer tells you is rejected. Every 小红书 item now carries a signed link plus the search-then-sign instruction. (6) n=75 could not be audited — there was no per-voice roster anywhere, so nobody could recount it. There is now a 65-row roster with a stated rule, and the number went DOWN, 75 → 66 → 65. No evidence was deleted; nine analysts and relays simply stopped being counted as voices. And one number in this round's own brief was wrong: we were told the Reddit thread stood at 187 and one commenter at 20; the live source read 183 and 19. We followed the source.
⚠️ ROUND 5 WAS THEN FAILED BY A THIRD INDEPENDENT VERIFIER, AND THIS IS WHAT HE FOUND — the worst single defect in this card's history. Four claims on this card face were FALSE ON THE DAY WE PUBLISHED THEM. We wrote, in the verdict, in a switch-if line, in a dropdown item and in the compatibility grid, that Fable 5 was included on no individual Claude plan at all. On 2026-07-20 — this card's own date — Anthropic made Fable 5 a permanent part of Max and Team Premium at 50% of usage limits, and Max is an individual plan starting at $100/month. Pro ($20) and Team Standard lost bundled access the same day and moved to usage credits at $10/M in, $50/M out after a one-time $100 credit; and the 50% slice is measured against a weekly baseline that also fell that day (~33%, the-decoder). It is corrected in all four places, and the corrected claim is sharper than the overreach: Fable 5 needs a $100+ Max plan, at half of a limit that just shrank. Here is the part that should cost us your trust rather than earn it back: the announcement was carried by two sources already on this card face. Simon Willison is evidence[0] — we quote his 07-16 K3 post and never opened his 07-18 post, which is the announcement. And 小红書 一起 Vibe is evidence_more[21], whose note is literally titled 「Fable 5 永不下线」 (“Fable 5 never goes offline”) — we quoted its last paragraph and never read its first, which opens 「Claude 果不其然怂了,宣布从7月20日起,Fable 5 永久在 Max 和 Team Premium 计划中,但限额仍然是50%」. Grepping all four receipt files for Team Premium and 永久在 Max returned zero hits. This was never weighed and rejected. It was never seen. A quote extracted from a page is not the same thing as a page that was read, and that is now the standing lesson of this card. The knock-on, stated plainly rather than buried: this card's strongest quantified finding — u/djdante's $20-vs-$20, 4× weekly quota — was measured 2026-07-15 against a Claude Pro tier that included Fable 5 and no longer does. It is not deleted: it was valid when made and a verifier re-read it verbatim at source. It is date-stamped everywhere it appears, as is u/ahnyudingslover's $20-tier corroboration and @m1nm13's “¥199 ≈ a Claude Pro Fable allowance”. And one more over-claim came out of the roster in the same pass: 小红書 老K was counted as a first-hand voice; his note is a 汇总, a compilation of Artificial Analysis data with no hands-on use, which the roster's own rule excludes. n: 66 → 65. Read our confidence with all of this in view: it stays LOW, and the reason is now sharper than it was — the defect landed on the subscription-economics half, the half this card calls “close to settled”. Near-settled describes a moving object. Check the plan tables yourself before you spend anything.

Show every source we read (111) Hide the ledger

The 111 sources in this card's ledger — 28 of them quoted above. Sources we read and did not quote are listed too, with why. This is the ledger, not a claim of exhaustiveness: anything the sweep read but deliberately left out of it is named in the coverage note above, with the reason. Reliability and neutrality are our own scores, not the publisher's. Vendor-owned pages are marked in the notes and never carry the verdict.

Stuck on a different switch?

If the card doesn't exist yet, request it — free, like everything here. Full sweep, weighted verdict, and one email the moment it's published.

Request a card