← Back to Blog
OperationsAug 27, 2026

Why Your Claude Costs Keep Climbing, and Three Questions to Start Fixing It

Why Your Claude Costs Keep Climbing, and Three Questions to Start Fixing It

A client of mine started using Claude Fable 5 for everything the day it launched: every draft, every quick question, every task that used to take her two minutes. Almost none of what she was doing needed Fable's level of capability. She'd just defaulted to the strongest model because it was there. That's exactly how you blast through your token budget for no real benefit. I want my clients getting the most value out of every dollar they spend on AI, and that starts with knowing which model actually earns its cost for the task in front of you.

Picking a model is both a quality and a cost decision, and the two aren't as linked as some people assume. Using the most powerful model for every task is like hiring a management consultant to answer your phones. It gets done well, but you're paying senior-partner rates for entry-level work.

The current lineup, and what I actually use each one for

Anthropic currently has five active Claude models, each built for a different point on the cost-versus-capability curve. Here's how I route real client work across them, and where each one tends to show up in everyday use.

Claude Haiku 4.5 is what I use for simple, high-volume tasks with one clear right answer (ex. “what is today's exchange rate?”), work with little room for ambiguity, fast and cheap. It's not usually a model I bring into client work on its own. It tends to show up as one piece of a larger project engagement, since most clients aren't coming to me specifically for the kind of task Haiku is built for.

When to use it:

  • Drafting first-pass replies to routine customer emails
  • Tagging or categorizing inbound leads before a call list is built
  • Quick summaries of long email threads or documents before deciding if they're worth reading in full
  • A single, unambiguous question with one clear matching answer, where you're not paying for reasoning you don't need

Claude Sonnet 5 is where most of my client work actually lives. It has enough reasoning depth for most judgment calls without the extra horsepower of Opus or Fable, so it's my default unless a task earns a step up, and that call gets sharper with experience.

From my own work:

  • Prepping raw email-thread data into a structured SharePoint list that became the knowledge base for a biotech client's support agent. Correctly separating multi-question emails and matching each part to the right entry mattered more than speed here, since the KB entries covered specific genetic test inquiries
  • A GEO/SEO audit tool that scores a website against defined criteria and produces a prioritized action plan: structured, rubric-driven analysis rather than open-ended judgment calls, which Sonnet handles cleanly
  • Synthesizing a client interview transcript into brand direction for a new website: the reasoning has real stakes, but the client and I are iterating together, so a strong first draft beats an over-engineered one
  • Supporting materials that repackage already-approved content, like turning a client's sales proposal into a follow-up deck: Sonnet's judgment on what to compress and what to carry over is exactly enough, with nothing new being generated

Other examples:

  • Drafting a client proposal or pitch from a rough outline
  • Restructuring a pricing sheet or service menu
  • Planning and executing a small research task (pull competitor info, organize the findings, draft a summary) without babysitting every step
  • General day-to-day writing and analysis where you're not sure which model to reach for

Claude Opus 4.8 is what I reserve for work where a mediocre output has a real, hard-to-reverse cost.

From my own work:

  • The live customer support agent for that same biotech client runs on Opus, searching the knowledge base and drafting the actual reply sent to the client: with health-adjacent, gene-specific information on the line, the accuracy of matching each question to the right answer matters more than the extra cost
  • I completed a competitive intelligence and voice-of-customer research project for a client, which meant synthesizing reviews and forum data across many sources into a ranked, prioritized report for the CEO. The entire value was in the quality of the synthesis and the judgment calls about what mattered
  • Generating a full sales proposal from a client discovery call transcript: that document directly affects whether my client can close their deals, so it gets the model built for depth
  • Governance work, like turning a board meeting transcript into official minutes and an action item register: the writing itself isn't complex, but it requires pre-trained judgment on what level of detail the Board expects.

Other examples:

  • Drafting a contract clause or terms you'll actually send to a client
  • Building a financial model with several interdependent assumptions
  • Untangling a messy dataset where a wrong conclusion costs real money
  • Any task where the honest question is “does this need this much horsepower, or am I just used to reaching for the strongest option”

Claude Fable 5 sits above Opus as the most capable model generally available. I'd reach for it on the rare document where the stakes and complexity both peak at once: an analysis with real financial and reputational consequences riding on it. For nearly everything else in a small business's day-to-day, it's the same overkill Opus can be relative to Sonnet, just at a steeper price.

Work I've done that I'd now route to Fable:

  • Series A and Series B financial modeling and investor presentations, where multiple interdependent assumptions all have to hold together and the reasoning directly shapes how much capital gets raised
  • Strategic and governance work tied to scaling an organization, where the stakes and the complexity of the reasoning are both high at once

Claude Mythos 5 is a variant of Fable 5 currently restricted to approved organizations under Anthropic's early-access program, so most businesses won't encounter it directly yet. I haven't at least.

Three questions to ask before you pick a model

Before starting any AI task, ask three questions:

  1. Is this high-volume and repetitive, or one-off and complex? Repetitive work belongs on the cheapest capable model; complex work can justify a stronger one.
  2. What's the cost of getting it wrong? A five-minute fix means use the cheaper model. A client-facing mistake means using tokens for the stronger one. The cost difference is trivial next to the cost of the error.
  3. Am I picking this out of habit, or because the task needs it? Most people default to whatever they used last time. Learning when to use which model extracts the most value out of what your business is already spending on AI.

As a starting rule of thumb: route routine, high-volume tasks to Haiku, run your default day-to-day work through Sonnet, and reserve Opus or Fable for the handful of tasks each week that are genuinely complex, high-stakes, or hard to undo if done poorly.