Cost Tracking for AI Requests: Monitor Token Costs

Setting Up Cost Tracking for AI Requests

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    918
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1032

Setting Up Cost Tracking for AI Requests

Imagine: Monday morning you see an OpenAI bill for $15,000, though you planned $10,000 a month. Sound familiar? Without tracking token costs by project and user, such overspend is only a matter of time. Cost tracking for LLMs is not an option but a necessity if you want to manage your AI infrastructure budget. Our experience shows: companies lose up to 40% of their budget on uncontrolled requests. We'll set up a system that calculates costs in real time for each model, project, and feature. As a result, you'll be able to plan expenses accurately and detect anomalies early.

What Problems Does Cost Tracking Solve?

Without an accounting system, you won't see which department or feature generates 80% of costs. A typical case: a RAG feature for customer support uses long contexts — 50% of all tokens go to it, but that's invisible until the bill arrives. Another problem is uncontrolled tests in dev environments: developers run heavy prompts on GPT-4o, though gpt-4o-mini would suffice for testing. Another pain point is the lack of alerts: if costs surge tenfold in a day due to a bug, you only learn about it at the end of the month. Our solution provides transparency: you see costs in real time, and when thresholds are exceeded — instant notification.

To understand the root cause, let's break down typical setup mistakes. First, many don't log the token count per request — without that, accurate cost calculation is impossible. Second, they ignore aggregation by feature: they only see the total bill, not what exactly is expensive. Third, they don't set budget limits — then any code error (e.g., an infinite loop calling the LLM) can burn through the monthly budget in an hour.

How We Set Up Token Accounting?

Each request to an LLM generates input and output tokens. The cost depends on the model and the number of tokens. For example, for GPT-4o the price per 1M input tokens is $2.50, output $10.00. For a local LLaMA 3-8B, costs are determined solely by GPU resources.

# Example pricing (update periodically via API) MODEL_PRICING = { "gpt-4o": {"input": 2.50, "output": 10.00}, "gpt-4o-mini": {"input": 0.15, "output": 0.60}, "claude-3-5-sonnet-20241022": {"input": 3.00, "output": 15.00}, "claude-3-haiku-20240307": {"input": 0.25, "output": 1.25}, "llama-3-8b-local": {"input": 0.0, "output": 0.0}, } def calculate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float: if model not in MODEL_PRICING: return 0.0 pricing = MODEL_PRICING[model] return (prompt_tokens * pricing["input"] + completion_tokens * pricing["output"]) / 1_000_000 

Aggregation by Dimensions

For analysis, we aggregate cost across multiple dimensions: project, user, feature tag (chat, RAG, classification), time interval. This allows answering the question: "Why did the AI bill increase by 30% in a week?"

class CostTracker: def record(self, request: LLMRequest, cost_usd: float): self.db.insert({ "timestamp": request.timestamp, "cost_usd": cost_usd, "model": request.model, "project_id": request.project_id, "user_id": request.user_id, "feature": request.feature_tag, "prompt_tokens": request.prompt_tokens, "completion_tokens": request.completion_tokens, }) def get_daily_by_project(self, days: int = 30) -> dict: return self.db.query(""" SELECT project_id, DATE(timestamp) as date, SUM(cost_usd) as total_cost, SUM(prompt_tokens) as total_tokens FROM llm_costs WHERE timestamp > NOW() - INTERVAL %s DAY GROUP BY project_id, date ORDER BY date, total_cost DESC """, (days,)) 

What Alerts Help Stay Within Budget?

Budget alerts are a key element of Cost Tracking. We configure two thresholds: daily and hourly. When 80% of the limit is reached, a warning is sent to Slack or Telegram; at 100%, automatic throttling of requests is triggered.

Threshold Action
80% daily limit Warning to team
100% daily limit Throttling (slow down) or blocking new requests
Anomaly >50% per day Notification with cause analysis
class BudgetGuard: def __init__(self, limits: dict): self.daily_limit_usd = limits["daily"] self.hourly_limit_usd = limits["hourly"] def check_budget(self, project_id: str) -> BudgetStatus: daily_spend = self.tracker.get_spend(project_id, hours=24) hourly_spend = self.tracker.get_spend(project_id, hours=1) alerts = [] if daily_spend > self.daily_limit_usd * 0.8: alerts.append(f"80% of daily budget consumed: ${daily_spend:.2f}/${self.daily_limit_usd:.2f}") if daily_spend > self.daily_limit_usd: alerts.append("DAILY BUDGET EXCEEDED — throttling enabled") return BudgetStatus(daily_spend=daily_spend, alerts=alerts, throttle_enabled=daily_spend > self.daily_limit_usd) 

Why Real-Time Monitoring Is Better?

Even an hour delay in accounting can lead to overspending. For example, launching a new feature without limits can cause costs to skyrocket tenfold in a day. A system with alerts reacts in minutes, not days. Our solution saves up to 30% of the budget through early anomaly detection.

What's Included in Cost Tracking Setup?

We provide:

  • Integration with provider logs (OpenAI, Anthropic, local models)
  • Dashboard with charts (daily cost, top features, cost per request)
  • Alert configuration in messengers
  • Documentation for adding new models
  • Support for one month after deployment

Implementation timeline: 5 to 15 days depending on infrastructure complexity. Cost is calculated individually — contact us for a project assessment. Get a consultation on AI cost optimization.

Comparison: Custom Tracker vs Our Solution

Criteria Custom Tracker Our Solution
Development time 2–4 weeks 5–15 days
Built-in alerts No Slack, Telegram, Email
Support for new models Manual API update
Accuracy guarantee No Unit tests and code review

Our solution reduces implementation time by 2-4 times compared to a custom tracker. Our team has over 5 years in ML infrastructure and 50+ Cost Tracking implementations for AI products. Contact us for a consultation.

Common Setup Mistakes - Not logging token count per request. - Ignoring aggregation by feature. - Not setting budget limits. - Using outdated model price lists. - Not testing alerts on real scenarios.