LlamaIndex
Auto-instrument LlamaIndex for RAG applications.
Risicare automatically instruments LlamaIndex for comprehensive RAG observability.
Installation
pip install 'risicare[llamaindex]'
# or
pip install risicare llama-indexVersion Compatibility
Requires llama-index-core >= 0.10.20.
The extra installs llama-index-core only. The samples on this page also use llama-index-llms-openai and llama-index-embeddings-openai. pip install llama-index installs both.
Auto-Instrumentation
import risicare
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
risicare.init()
documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
# Automatically traced
response = query_engine.query("What is Risicare?")What's Captured
| Feature | Description |
|---|---|
| Query Engine | Full query execution |
| Retrievers | Document retrieval. The retrieved nodes, with their scores, are in the span text when content capture is on |
| LLM Calls | Synthesis and completion calls. A non-streaming call gets no provider span; a streamed call also gets one |
| Embeddings | Embedding generation |
| Node Parsing | Document chunking |
| Index Operations | Index building and updates |
Span Names
llamaindex.query/{cls}, llamaindex.retrieve/{cls}, llamaindex.embedding/{model},
llamaindex.synthesize/{cls}, llamaindex.llm/{model} and llamaindex.component/{cls}.
One run is one trace (Python 0.5.1 and later)
Since Python 0.5.1, the spans of one run are one trace: a span takes its trace and parent
from its LlamaIndex parent span, and a span with no LlamaIndex parent joins the active Risicare
span (a trace, an agent, a session). In 0.5.0 each LlamaIndex span was a trace of its own.
A streamed LLM call (streaming=True, stream_complete()) still gets a provider span next
to the llamaindex.llm span. The two spans are in one trace only when the call runs inside
a Risicare span; if not, the provider span is a trace of its own. Because LlamaIndex nests its
own calls, one query() can send several query, synthesize and llm spans for one model
call. Measured with llama-index-core 0.14.25 and Python 0.6.0.
Provider Deduplication
Provider Deduplication
With the Python SDK, the provider span of a non-streaming LLM or embedding call is suppressed. A streamed LLM call (streaming=True, stream_complete()) also sends the provider span (see the notice above). Other provider span types (e.g., tool calls) are not suppressed. For JavaScript, see JavaScript / TypeScript below.
Query Engines
All query engine types are traced:
# Simple query engine
query_engine = index.as_query_engine()
# Chat engine
chat_engine = index.as_chat_engine()
response = chat_engine.chat("Hello!")
# Retriever-based
retriever = index.as_retriever()
nodes = retriever.retrieve("query")Retrievers
Retrieval operations capture document details:
retriever = index.as_retriever(similarity_top_k=5)
nodes = retriever.retrieve("What is AI?")
# With content capture on, the span text holds each node and its scoreAgents
A LlamaIndex agent sends spans for its components, its LLM calls and its tools. The spans of one run are one trace, with the agent span (llamaindex.component/ReActAgent for the sample below) as the root. In this sample the agent calls the LLM with a streamed call, so each LLM call also gets a provider span. That span is a trace of its own unless the run is inside a Risicare span (see the notice above).
import asyncio
from llama_index.core.agent.workflow import ReActAgent
from llama_index.core.tools import FunctionTool
from llama_index.llms.openai import OpenAI
llm = OpenAI(model="gpt-4o")
def search(query: str) -> str:
"""Search for information."""
return f"Results for {query}"
tool = FunctionTool.from_defaults(fn=search)
agent = ReActAgent(tools=[tool], llm=llm)
async def main():
response = await agent.run("Search for AI news")
print(response)
asyncio.run(main())The workflow agent of this sample is in llama-index-core 0.12 and later. ReActAgent.from_tools() and agent.chat() are the older API: they run on llama-index-core 0.12, and 0.13 and later do not have them.
Embeddings
Embedding operations are captured:
from llama_index.embeddings.openai import OpenAIEmbedding
embed_model = OpenAIEmbedding()
embedding = embed_model.get_text_embedding("Hello")
# The embedding span records the model nameIndex Building
Index creation is traced:
# Document loading
documents = SimpleDirectoryReader("data").load_data()
# Index building (chunking + embedding)
index = VectorStoreIndex.from_documents(
documents,
show_progress=True
)
# Node parsing and embedding spansStreaming
query_engine = index.as_query_engine(streaming=True)
response = query_engine.query("Write about AI")
for text in response.response_gen:
print(text, end="")JavaScript / TypeScript
The JS SDK provides RisicareLlamaIndexHandler for LlamaIndex.TS:
import { Settings } from 'llamaindex';
import { RisicareLlamaIndexHandler } from 'risicare/llamaindex';
const handler = new RisicareLlamaIndexHandler();
for (const name of RisicareLlamaIndexHandler.EVENT_NAMES) {
Settings.callbackManager.on(name as any, (e: any) => handler.onEvent(e));
}
// Provider spans are suppressed only inside withSuppression():
const result = await handler.withSuppression(() => queryEngine.query({ query: '...' }));The handler does not suppress provider instrumentation by itself. If you also patch the provider client, run your calls inside handler.withSuppression() to prevent duplicate spans.