Cost Tracking
Track LLM costs across providers.
Risicare automatically tracks LLM costs across 14 providers using a built-in per-model pricing table. The LLM Cost KPI card on the main dashboard shows total spend, and the Models chart breaks down requests by model — visible in the dashboard overview.
Automatic Cost Calculation
The SDK captures token counts on every LLM call. The server calculates the dollar cost from those tokens and its pricing table. You don't compute cost yourself.
The server prices a span on its model id: the response model (gen_ai.response.model), or else the request model (gen_ai.request.model).
import risicare
from openai import OpenAI
risicare.init()
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}]
)
# The SDK captures token counts:
# - gen_ai.usage.prompt_tokens: 10
# - gen_ai.usage.completion_tokens: 15
#
# Risicare then computes the cost server-side and stores it as:
# - llm_cost_usd: 0.000175 (stored cost column)How Risicare calculates cost
The server's price table wins. For a model that the table knows, the dashboard shows the server's cost when it is above zero, also when the span carries a cost of its own. The JavaScript SDK (up to and including 0.9.0) sends a cost from its own older table, and a span sent over OTLP can carry gen_ai.usage.cost: the server replaces both for a known model. For a model that the table does not know, or when the server's cost is zero (for example, no token counts), a cost that the span carries is kept.
Matching. The model id of the span must equal an entry of the table or one of its aliases. The comparison ignores case, and there is no prefix match. The table has Anthropic's list prices for the current Claude models (read on 2026-10-04). A span whose model is the bare name opus, sonnet, haiku or fable is priced as the model that Claude Code uses for that name on the Anthropic API: Claude Opus 5.5, Sonnet 5.5, Haiku 4.5 and Fable 5.1. On Amazon Bedrock, Google Cloud and Microsoft Foundry, Claude Code can use an older model for sonnet or opus, and then this cost is wrong.
Labels in the dashboard:
- On a span (the span detail): the cost of a model that the table knows has no label. A cost at the fallback price ($0.15 input and $0.60 output per 1M tokens, for a model that the table does not know and a span that carries no cost) is marked "(est.)". A span with no cost shows no cost.
- On a total (the LLM Cost card, and the Total Cost cards of the Sessions and Agents pages): "estimated — no price table for this model" when each priced span is at the fallback price; "partly estimated" (with the estimated amount) when some are; "no cost basis recorded" when no span has a cost.
Not priced yet:
- Cached tokens. See Cache Token Support: for Anthropic models the cost of cached traffic is too low.
- Some model-name forms get the fallback price: an Amazon Bedrock id of a Claude model newer than Claude 3.5 (for example
anthropic.claude-sonnet-5-5orus.anthropic.claude-…-v1:0, and ARNs), a dated Google Vertex id (claude-…@20251101), and a model name in theprovider/modelform (for example from LiteLLM or OpenRouter). - Price tiers. Batch discounts, fast or priority tiers, regional or data-residency premiums and long-context tiers are not applied. The cost is at the standard list price.
- Some current models of other providers are not in the table yet (for example newer OpenAI and Gemini models). If the span carries no cost, they get the fallback price and the "estimated" label. If it carries a cost (for example from the JavaScript SDK), that cost is shown with no label.
Supported Providers
| Provider | Pricing | Cached Rate* |
|---|---|---|
| OpenAI | Per-token | 50% cached discount |
| Anthropic | Per-token | 90% to 97.5% cached discount, by model |
| Per-token | - | |
| Cohere | Per-token | - |
| Mistral | Per-token | - |
| Groq | Per-token | - |
| Together AI | Per-token | - |
| Amazon Bedrock | Per-token | - |
| Vertex AI | Per-token | - |
| Cerebras | Per-token | - |
| HuggingFace | Per-token | - |
| Fireworks | Per-token | - |
| xAI | Per-token | - |
| Ollama | Free for the local models in the pricing table (a total of only such spans shows "no cost basis recorded") | - |
* The cached rate is part of the pricing model but is not yet applied end-to-end — see Cache Token Support below.
Pricing Examples
The cost in the dashboard comes from the server's price table. It does not come from your invoice. The tables below show some of the rates in the server table (per 1M tokens).
OpenAI
| Model | Input | Output | Cached Input |
|---|---|---|---|
| gpt-4o | $2.50 | $10.00 | $1.25 |
| gpt-4o-mini | $0.15 | $0.60 | $0.075 |
| o1 | $15.00 | $60.00 | $7.50 |
| o1-mini | $3.00 | $12.00 | $1.50 |
| gpt-4-turbo | $10.00 | $30.00 | - |
Anthropic
| Model | Input | Output | Cached Input |
|---|---|---|---|
| claude-fable-5-1 | $10.00 | $50.00 | $0.25 |
| claude-opus-5-5 | $4.00 | $20.00 | $0.20 |
| claude-sonnet-5-5 | $2.00 | $10.00 | $0.20 |
| claude-opus-4-5-20251101 | $5.00 | $25.00 | $0.50 |
| claude-sonnet-4-5-20250929 | $3.00 | $15.00 | $0.30 |
| claude-haiku-4-5-20251001 | $1.00 | $5.00 | $0.10 |
| Model | Input | Output |
|---|---|---|
| gemini-2.0-pro | $1.25 | $5.00 |
| gemini-2.0-flash | $0.10 | $0.40 |
| gemini-1.5-pro | $1.25 | $5.00 |
| gemini-1.5-flash | $0.075 | $0.30 |
Cache Token Support
Cached-token discounting is not applied
The pricing table includes cached-input rates (the Cached Input columns above), but the SDKs do not send cached-token counts, and the server prices only the prompt and completion tokens that it receives. So prompt-caching discounts are not reflected in llm_cost_usd:
- For OpenAI the reported cost is an upper bound: the cached tokens are part of the prompt tokens, and all of them are priced at the full input rate.
- For Anthropic, and for Claude on AWS Bedrock and Google Vertex, the cost of cached traffic is too low: the input count leaves out cache reads and cache writes, so those tokens are not priced.
A fix is planned (server and SDK).
Provider-side prompt caching still reduces your provider bill. The cost that Risicare reports does not reflect it.
# Provider-side caching works as usual and reduces your real provider bill:
response = anthropic.messages.create(
model="claude-sonnet-5-5",
system=[{
"type": "text",
"text": long_prompt,
"cache_control": {"type": "ephemeral"}
}],
messages=[...]
)
# Note: the cost that Risicare stores does not include the cached tokens of
# this call.Dashboard Views
The dashboard shows cost in these places:
- The LLM Cost card on the project dashboard: the total for the selected time range. The card says when part of the total is an estimate.
- The trace list and the trace page: the cost of each trace.
- The Sessions page: the cost of each session, and a Total Cost card that says when part of it is an estimate.
- The Agents page: a Total Cost card, which says when part of it is an estimate.
- The agent page and the Top Performing Agents card of the project dashboard: the cost of each agent.
- The span detail: the cost of one LLM call. An estimate from the fallback rate is marked "(est.)".
There is no cost view by provider, by model or by feature. The Models card shows request counts by model.
API Access
Not available with an API key during the beta: the Management API cannot be reached, so you cannot read cost data over HTTP. Use the dashboard for the cost of a trace, a session or an agent, and for the project's total LLM cost over the selected time range.
Cost Alerts
Not available during the beta
During the beta the Management API is not available with an API key, and the
dashboard's Alerts page has no control that creates a rule. Alert
evaluation is turned off during the beta: no rule fires, and no
alert.triggered event is sent.
As designed, an alert rule fires when spend over a window crosses a threshold. A
rule has these fields; name, metric, operator, and threshold are required.
| Field | Required | Notes |
|---|---|---|
name | yes | Human-readable rule name |
description | no | Free-text description |
metric | yes | Must be one of the supported metrics below |
operator | yes | gt, lt, eq, gte, lte, ne |
threshold | yes | Numeric threshold |
window_minutes | no | Evaluation window, default 5 |
severity | no | info, warning, error, critical — default warning |
Supported metrics. A rule can name only these metrics:
error_rate · latency_p50 · latency_p95 · latency_p99 · trace_volume ·
cost_usd · total_tokens
cost_usd is measured over the rule's own window (window_minutes), not over a
calendar day.
Webhook is the only alert channel, and it is held
As designed, when a rule fires, the alert worker emits an alert.triggered
webhook event to every webhook in the project subscribed to it — that is the
only delivery path there is.
During the beta no rule fires, and the dashboard has no control that registers a webhook (see Webhooks). Cost alerting does not work during the beta.
Cost Optimization
When the diagnosis pipeline analyzes a budget- or resource-related error, it can surface cost-oriented fix recommendations, such as a "use a cheaper model when approaching budget" suggestion. These appear as recommendations in the diagnosis output — they are not auto-applied. Diagnosis is held for the beta, so no new recommendations of this kind are produced today, and fix deployment is gated off (see Self-Healing).
- Model downgrade suggestions: e.g. "Switch to a cheaper model when approaching the budget limit"
- Caching / token-reduction hints surfaced alongside a diagnosis, where applicable
Risicare does not currently run a standalone, always-on cost-optimization analysis independent of the diagnosis pipeline.