Skip to main content
GitHub

LlamaIndex

Auto-instrument LlamaIndex for RAG applications.

Risicare automatically instruments LlamaIndex for comprehensive RAG observability.

Installation

pip install 'risicare[llamaindex]'
# or
pip install risicare llama-index

Version Compatibility

Requires llama-index-core >= 0.10.20.

The extra installs llama-index-core only. The samples on this page also use llama-index-llms-openai and llama-index-embeddings-openai. pip install llama-index installs both.

Auto-Instrumentation

import risicare
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
 
risicare.init()
 
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
 
# Automatically traced
response = query_engine.query("What is Risicare?")

What's Captured

FeatureDescription
Query EngineFull query execution
RetrieversDocument retrieval. The retrieved nodes, with their scores, are in the span text when content capture is on
LLM CallsSynthesis and completion calls. A non-streaming call gets no provider span; a streamed call also gets one
EmbeddingsEmbedding generation
Node ParsingDocument chunking
Index OperationsIndex building and updates

Span Names

llamaindex.query/{cls}, llamaindex.retrieve/{cls}, llamaindex.embedding/{model}, llamaindex.synthesize/{cls}, llamaindex.llm/{model} and llamaindex.component/{cls}.

One run is one trace (Python 0.5.1 and later)

Since Python 0.5.1, the spans of one run are one trace: a span takes its trace and parent from its LlamaIndex parent span, and a span with no LlamaIndex parent joins the active Risicare span (a trace, an agent, a session). In 0.5.0 each LlamaIndex span was a trace of its own. A streamed LLM call (streaming=True, stream_complete()) still gets a provider span next to the llamaindex.llm span. The two spans are in one trace only when the call runs inside a Risicare span; if not, the provider span is a trace of its own. Because LlamaIndex nests its own calls, one query() can send several query, synthesize and llm spans for one model call. Measured with llama-index-core 0.14.25 and Python 0.6.0.

Provider Deduplication

Provider Deduplication

With the Python SDK, the provider span of a non-streaming LLM or embedding call is suppressed. A streamed LLM call (streaming=True, stream_complete()) also sends the provider span (see the notice above). Other provider span types (e.g., tool calls) are not suppressed. For JavaScript, see JavaScript / TypeScript below.

Query Engines

All query engine types are traced:

# Simple query engine
query_engine = index.as_query_engine()
 
# Chat engine
chat_engine = index.as_chat_engine()
response = chat_engine.chat("Hello!")
 
# Retriever-based
retriever = index.as_retriever()
nodes = retriever.retrieve("query")

Retrievers

Retrieval operations capture document details:

retriever = index.as_retriever(similarity_top_k=5)
nodes = retriever.retrieve("What is AI?")
 
# With content capture on, the span text holds each node and its score

Agents

A LlamaIndex agent sends spans for its components, its LLM calls and its tools. The spans of one run are one trace, with the agent span (llamaindex.component/ReActAgent for the sample below) as the root. In this sample the agent calls the LLM with a streamed call, so each LLM call also gets a provider span. That span is a trace of its own unless the run is inside a Risicare span (see the notice above).

import asyncio
from llama_index.core.agent.workflow import ReActAgent
from llama_index.core.tools import FunctionTool
from llama_index.llms.openai import OpenAI
 
llm = OpenAI(model="gpt-4o")
 
def search(query: str) -> str:
    """Search for information."""
    return f"Results for {query}"
 
tool = FunctionTool.from_defaults(fn=search)
agent = ReActAgent(tools=[tool], llm=llm)
 
async def main():
    response = await agent.run("Search for AI news")
    print(response)
 
asyncio.run(main())

The workflow agent of this sample is in llama-index-core 0.12 and later. ReActAgent.from_tools() and agent.chat() are the older API: they run on llama-index-core 0.12, and 0.13 and later do not have them.

Embeddings

Embedding operations are captured:

from llama_index.embeddings.openai import OpenAIEmbedding
 
embed_model = OpenAIEmbedding()
embedding = embed_model.get_text_embedding("Hello")
 
# The embedding span records the model name

Index Building

Index creation is traced:

# Document loading
documents = SimpleDirectoryReader("data").load_data()
 
# Index building (chunking + embedding)
index = VectorStoreIndex.from_documents(
    documents,
    show_progress=True
)
# Node parsing and embedding spans

Streaming

query_engine = index.as_query_engine(streaming=True)
response = query_engine.query("Write about AI")
 
for text in response.response_gen:
    print(text, end="")

JavaScript / TypeScript

The JS SDK provides RisicareLlamaIndexHandler for LlamaIndex.TS:

import { Settings } from 'llamaindex';
import { RisicareLlamaIndexHandler } from 'risicare/llamaindex';
 
const handler = new RisicareLlamaIndexHandler();
for (const name of RisicareLlamaIndexHandler.EVENT_NAMES) {
  Settings.callbackManager.on(name as any, (e: any) => handler.onEvent(e));
}
 
// Provider spans are suppressed only inside withSuppression():
const result = await handler.withSuppression(() => queryEngine.query({ query: '...' }));

The handler does not suppress provider instrumentation by itself. If you also patch the provider client, run your calls inside handler.withSuppression() to prevent duplicate spans.

Next Steps