Setting Up Cost Tracking for AI Requests
Imagine: Monday morning you see an OpenAI bill for $15,000, though you planned $10,000 a month. Sound familiar? Without tracking token costs by project and user, such overspend is only a matter of time. Cost tracking for LLMs is not an option but a necessity if you want to manage your AI infrastructure budget. Our experience shows: companies lose up to 40% of their budget on uncontrolled requests. We'll set up a system that calculates costs in real time for each model, project, and feature. As a result, you'll be able to plan expenses accurately and detect anomalies early.
What Problems Does Cost Tracking Solve?
Without an accounting system, you won't see which department or feature generates 80% of costs. A typical case: a RAG feature for customer support uses long contexts — 50% of all tokens go to it, but that's invisible until the bill arrives. Another problem is uncontrolled tests in dev environments: developers run heavy prompts on GPT-4o, though gpt-4o-mini would suffice for testing. Another pain point is the lack of alerts: if costs surge tenfold in a day due to a bug, you only learn about it at the end of the month. Our solution provides transparency: you see costs in real time, and when thresholds are exceeded — instant notification.
To understand the root cause, let's break down typical setup mistakes. First, many don't log the token count per request — without that, accurate cost calculation is impossible. Second, they ignore aggregation by feature: they only see the total bill, not what exactly is expensive. Third, they don't set budget limits — then any code error (e.g., an infinite loop calling the LLM) can burn through the monthly budget in an hour.
How We Set Up Token Accounting?
Each request to an LLM generates input and output tokens. The cost depends on the model and the number of tokens. For example, for GPT-4o the price per 1M input tokens is $2.50, output $10.00. For a local LLaMA 3-8B, costs are determined solely by GPU resources.
# Example pricing (update periodically via API) MODEL_PRICING = { "gpt-4o": {"input": 2.50, "output": 10.00}, "gpt-4o-mini": {"input": 0.15, "output": 0.60}, "claude-3-5-sonnet-20241022": {"input": 3.00, "output": 15.00}, "claude-3-haiku-20240307": {"input": 0.25, "output": 1.25}, "llama-3-8b-local": {"input": 0.0, "output": 0.0}, } def calculate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float: if model not in MODEL_PRICING: return 0.0 pricing = MODEL_PRICING[model] return (prompt_tokens * pricing["input"] + completion_tokens * pricing["output"]) / 1_000_000 Aggregation by Dimensions
For analysis, we aggregate cost across multiple dimensions: project, user, feature tag (chat, RAG, classification), time interval. This allows answering the question: "Why did the AI bill increase by 30% in a week?"
class CostTracker: def record(self, request: LLMRequest, cost_usd: float): self.db.insert({ "timestamp": request.timestamp, "cost_usd": cost_usd, "model": request.model, "project_id": request.project_id, "user_id": request.user_id, "feature": request.feature_tag, "prompt_tokens": request.prompt_tokens, "completion_tokens": request.completion_tokens, }) def get_daily_by_project(self, days: int = 30) -> dict: return self.db.query(""" SELECT project_id, DATE(timestamp) as date, SUM(cost_usd) as total_cost, SUM(prompt_tokens) as total_tokens FROM llm_costs WHERE timestamp > NOW() - INTERVAL %s DAY GROUP BY project_id, date ORDER BY date, total_cost DESC """, (days,)) What Alerts Help Stay Within Budget?
Budget alerts are a key element of Cost Tracking. We configure two thresholds: daily and hourly. When 80% of the limit is reached, a warning is sent to Slack or Telegram; at 100%, automatic throttling of requests is triggered.
| Threshold | Action |
|---|---|
| 80% daily limit | Warning to team |
| 100% daily limit | Throttling (slow down) or blocking new requests |
| Anomaly >50% per day | Notification with cause analysis |
class BudgetGuard: def __init__(self, limits: dict): self.daily_limit_usd = limits["daily"] self.hourly_limit_usd = limits["hourly"] def check_budget(self, project_id: str) -> BudgetStatus: daily_spend = self.tracker.get_spend(project_id, hours=24) hourly_spend = self.tracker.get_spend(project_id, hours=1) alerts = [] if daily_spend > self.daily_limit_usd * 0.8: alerts.append(f"80% of daily budget consumed: ${daily_spend:.2f}/${self.daily_limit_usd:.2f}") if daily_spend > self.daily_limit_usd: alerts.append("DAILY BUDGET EXCEEDED — throttling enabled") return BudgetStatus(daily_spend=daily_spend, alerts=alerts, throttle_enabled=daily_spend > self.daily_limit_usd) Why Real-Time Monitoring Is Better?
Even an hour delay in accounting can lead to overspending. For example, launching a new feature without limits can cause costs to skyrocket tenfold in a day. A system with alerts reacts in minutes, not days. Our solution saves up to 30% of the budget through early anomaly detection.
What's Included in Cost Tracking Setup?
We provide:
- Integration with provider logs (OpenAI, Anthropic, local models)
- Dashboard with charts (daily cost, top features, cost per request)
- Alert configuration in messengers
- Documentation for adding new models
- Support for one month after deployment
Implementation timeline: 5 to 15 days depending on infrastructure complexity. Cost is calculated individually — contact us for a project assessment. Get a consultation on AI cost optimization.
Comparison: Custom Tracker vs Our Solution
| Criteria | Custom Tracker | Our Solution |
|---|---|---|
| Development time | 2–4 weeks | 5–15 days |
| Built-in alerts | No | Slack, Telegram, Email |
| Support for new models | Manual | API update |
| Accuracy guarantee | No | Unit tests and code review |
Our solution reduces implementation time by 2-4 times compared to a custom tracker. Our team has over 5 years in ML infrastructure and 50+ Cost Tracking implementations for AI products. Contact us for a consultation.







