Cost Control for LLMs: Token Counting and Budgeting in Production

How Not to Lose Your Budget on LLMs: Token Counting and Budgeting

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

How Not to Lose Your Budget on LLMs: Token Counting and Budgeting

Picture this: a development team integrates GPT-4 for processing customer requests, and a month later the API bill hits $5,000 instead of the expected $500. The reason—every request included a 2,000-token system prompt, and duplicate requests were never cached. Token counting and budgeting are not optional; they are a necessity for any production system using LLMs. Without them, you risk a bill 5–10 times higher than anticipated.

We implemented a cost management system for a SaaS platform handling 10,000 daily requests to GPT-4 and Claude. The result: a 35% cost reduction in the first month, saving over $1,500 per month. Below we break down how we did it and how you can replicate the success. Our certified MLOps engineers guarantee a minimum 20% cost reduction for most clients, backed by a proven track record with over 50 successful projects.

Why Budgeting LLM Requests Matters

LLM request costs are unpredictable: output token length varies, and system prompts often bloat unchecked. Without budgeting, it's easy to get a bill 5–10 times higher than expected. Budgeting solves three problems:

  • Preventing overspend: alerts when daily or monthly limits are exceeded.
  • Cost optimization: identifying inefficient prompts and models.
  • Transparency for teams: dashboards with breakdowns by user, feature, and model.

Principles of Token Counting in Production

Token counting estimates the number of tokens in a request before sending. For OpenAI we use the tiktoken library (OpenAI Tokenizer documentation), for Anthropic the built-in counter. Example calculation:

import tiktoken def count_tokens_openai(text: str, model: str = "gpt-4") -> int: enc = tiktoken.encoding_for_model(model) return len(enc.encode(text)) def estimate_request_cost(prompt: str, max_completion: int = 1000, model: str = "gpt-4-turbo") -> dict: input_tokens = count_tokens_openai(prompt, model) total_tokens = input_tokens + max_completion prices = { "gpt-4-turbo": {"input": 10.0, "output": 30.0}, "gpt-4o": {"input": 5.0, "output": 15.0}, "gpt-4o-mini": {"input": 0.15, "output": 0.60}, "claude-3-5-sonnet": {"input": 3.0, "output": 15.0}, } price = prices.get(model, {"input": 10.0, "output": 30.0}) estimated_cost = (input_tokens / 1_000_000 * price["input"] + max_completion / 1_000_000 * price["output"]) return {"input_tokens": input_tokens, "max_output_tokens": max_completion, "estimated_cost_usd": estimated_cost} 

Cost estimation allows rejecting requests that exceed a user's or team's limit. The tokenizer vocabulary and encoding method are critical for accurate counting.

Model Cost Comparison Table

Model Input per 1M tokens Output per 1M tokens
GPT-4-turbo $10.00 $30.00
GPT-4o $5.00 $15.00
GPT-4o-mini $0.15 $0.60
Claude 3.5 Sonnet $3.00 $15.00

Methods to Reduce LLM Costs

One effective method is automatic model routing. For example, for simple tasks (data extraction, basic classification) we use GPT-4o-mini, which is 10x cheaper than GPT-4-turbo yet yields comparable results. For complex generations we keep the top-tier models. In our case study, this reduced the average request cost from $0.03 to $0.008.

We also implement caching: duplicate requests (e.g., repeated user questions) are served from Redis cache, saving up to 40% of the budget. We use an LRU cache with a 24-hour TTL. System prompt optimization can further cut input tokens by half, significantly reducing API cost management overhead.

Budgeting System with Redis and Middleware

We use Redis to store limits—it's fast and supports atomic operations. Our Redis budgeting system stores daily and monthly limits per user and globally. Budget structure:

from dataclasses import dataclass, field import threading @dataclass class TokenBudget: daily_limit_usd: float monthly_limit_usd: float per_user_daily_limit_usd: float = 1.0 spent_today: float = field(default=0.0) spent_month: float = field(default=0.0) _lock: threading.Lock = field(default_factory=threading.Lock) class LLMBudgetManager: def __init__(self, redis_client, budget: TokenBudget): self.redis = redis_client self.budget = budget def check_and_reserve(self, user_id: str, estimated_cost: float) -> bool: """Check budget before request""" daily_spent = float(self.redis.get(f"budget:daily") or 0) if daily_spent + estimated_cost > self.budget.daily_limit_usd: raise BudgetExceededError(f"Daily budget ${self.budget.daily_limit_usd} exceeded") user_spent = float(self.redis.get(f"budget:user:{user_id}:daily") or 0) if user_spent + estimated_cost > self.budget.per_user_daily_limit_usd: raise BudgetExceededError(f"User daily budget ${self.budget.per_user_daily_limit_usd} exceeded") pipe = self.redis.pipeline() pipe.incrbyfloat(f"budget:daily", estimated_cost) pipe.expire(f"budget:daily", 86400) pipe.incrbyfloat(f"budget:user:{user_id}:daily", estimated_cost) pipe.expire(f"budget:user:{user_id}:daily", 86400) pipe.execute() return True def record_actual_cost(self, user_id: str, actual_cost: float, estimated_cost: float): correction = actual_cost - estimated_cost if abs(correction) > 0.001: self.redis.incrbyfloat("budget:daily", correction) 
Middleware for automatic cost tracking
import functools def track_llm_cost(model: str = "gpt-4o"): def decorator(func): @functools.wraps(func) async def wrapper(*args, **kwargs): result = await func(*args, **kwargs) if hasattr(result, 'usage'): cost = compute_cost(result.usage.prompt_tokens, result.usage.completion_tokens, model) analytics.record(function=func.__name__, model=model, cost=cost, input_tokens=result.usage.prompt_tokens, output_tokens=result.usage.completion_tokens) return result return wrapper return decorator 

Typical result of implementation: 20–40% reduction in LLM costs by identifying inefficient requests (long unnecessary system prompts, duplicate uncached requests) and choosing the right model for each task. Budget alerts via Slack or Telegram ensure no surprises.

Practical Results of Token Counting

Without token counting, you don't know how much you spend per request. We implemented dashboards in Grafana showing cost by user, model, and function. This allowed the client to discover that 30% of queries were repeats with the same prompts. After enabling caching, savings reached $1,500 per month. LLM model selection and system prompt optimization further cut costs. Want similar savings? Contact us for a cost audit.

Common Mistakes in LLM Budgeting

We often see the same mistakes: ignoring system prompt tokens, lack of caching, using expensive models for simple tasks, no per-user limits, and no monitoring. Each can increase costs 10–50 times. The solution: implement token counting, budget alerts, and model routing. Our Redis budgeting system also tracks per-user limits to prevent abuse.

Step-by-Step Token Counting Setup

  1. Integrate the tiktoken library.
  2. Implement a token counting function for your model.
  3. Estimate request cost using model pricing (see table above).
  4. Add budget check before API calls.
  5. Set up logging and dashboards.

Scope of Work for Budgeting System Implementation

We offer turnkey implementation of token counting and budgeting systems. The project includes:

  • Audit of current architecture and request profile.
  • Integration of tiktoken and Anthropic SDK into your codebase.
  • Redis setup for storing limits and statistics.
  • Implementation of middleware for automatic cost tracking.
  • Monitoring dashboards (Grafana + Prometheus).
  • Team training and documentation.

Our MLOps experience spans over 5 years; we have implemented budgeting systems for companies handling up to 1 million requests per day, serving 20+ enterprise clients with a proven methodology. Implementation timelines: 2–4 weeks for a basic system, up to 6 weeks for complex architectures. Contact us for a project assessment and request an LLM cost audit. We guarantee a minimum 20% cost reduction or your money back.