Skip to main content
GitHub

LiteLLM

Auto-instrument LiteLLM for unified LLM access.

Risicare automatically instruments LiteLLM for unified access to 100+ LLM providers.

Python only

This framework integration is available in the Python SDK only. No JavaScript package exists for LiteLLM.

Installation

pip install 'risicare[litellm]'
# or
pip install risicare litellm

Version Compatibility

Requires litellm >= 1.30.0.

Auto-Instrumentation

import risicare
import litellm
 
risicare.init()
 
# Automatically traced
response = litellm.completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

What's Captured

FeatureDescription
Completion CallsAll completion/acompletion calls
Provider RoutingModel-to-provider mapping
FallbacksFallback chain execution
Cost TrackingLiteLLM's own cost, in the span attribute gen_ai.usage.cost_usd
Load BalancingRouter selections

Span Hierarchy

litellm.completion/{model} (LLM_CALL kind)

Provider Deduplication

Provider Deduplication

When using LiteLLM, underlying LLM provider spans are automatically suppressed to avoid duplicate traces. You don't need to disable provider instrumentation manually.

Multiple Providers

LiteLLM's unified interface works with all providers:

# OpenAI
response = litellm.completion(model="gpt-4o", messages=[...])
 
# Anthropic
response = litellm.completion(model="claude-sonnet-5-5", messages=[...])
 
# Bedrock
response = litellm.completion(model="bedrock/claude-3-sonnet", messages=[...])
 
# Together AI
response = litellm.completion(model="together_ai/meta-llama/Llama-3-70b", messages=[...])

Streaming

response = litellm.completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Write a story"}],
    stream=True
)
 
for chunk in response:
    print(chunk.choices[0].delta.content or "", end="")

Fallbacks

Fallback chains are fully traced:

response = litellm.completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}],
    fallbacks=["claude-sonnet-5-5", "gemini/gemini-2.5-pro"]
)
 
# Each fallback attempt appears as a child span

Router

The LiteLLM Router is instrumented:

from litellm import Router
 
router = Router(
    model_list=[
        {"model_name": "gpt-4", "litellm_params": {"model": "gpt-4o"}},
        {"model_name": "gpt-4", "litellm_params": {"model": "azure/gpt-4"}},
    ]
)
 
# Load balancing decisions are captured
response = router.completion(model="gpt-4", messages=[...])

Embeddings

response = litellm.embedding(
    model="text-embedding-ada-002",
    input=["Hello, world!"]
)

Cost Tracking

LiteLLM's own cost calculation is recorded in the span attribute gen_ai.usage.cost_usd. The cost that the dashboard shows does not use it: it comes from the server's price table (see Cost Tracking). If the model name on the span is in the provider/model form, it gets the fallback price there.

from litellm import completion_cost
 
response = litellm.completion(model="gpt-4o", messages=[...])
cost = completion_cost(completion_response=response)
 
# LiteLLM's cost appears in the span attribute gen_ai.usage.cost_usd

Next Steps