Core Concepts
Understand the key concepts in Risicare observability.
This guide explains the core concepts you'll encounter when using Risicare.
Traces and Spans
Traces
A trace represents a complete execution flow through your agent. It starts when your agent receives input and ends when it produces output.
Trace: "Answer user question about weather"
├── Span: parse_input (2ms)
├── Span: llm_call to gpt-4o (1.2s)
├── Span: tool_call to weather_api (300ms)
└── Span: format_response (5ms)
Spans
A span represents a single unit of work within a trace. Spans have:
- Name: What operation this span represents
- Kind: The type of span (LLM_CALL, TOOL_CALL, AGENT, etc.)
- Timing: Start time and duration
- Attributes: Key-value metadata
- Events: Named events that occurred during the span
- Status: UNSET (the default), OK or ERROR
Span Hierarchy
Spans form a tree structure with parent-child relationships:
Create nested spans with the tracer returned by get_tracer(). Each
start_span() is a context manager, and a span opened inside another is
automatically parented to it:
from risicare import get_tracer, SpanKind
tracer = get_tracer()
with tracer.start_span("process_request") as parent:
with tracer.start_span("call_llm", kind=SpanKind.LLM_CALL) as child:
# child.parent_span_id == parent.span_id
# child.trace_id == parent.trace_id
passThere is no top-level start_span()
start_span is a method on the tracer, not a module-level export —
from risicare import start_span raises ImportError. Always go through
get_tracer(). Spans are only recorded after risicare.init(); before
that the tracer returns inert no-op spans.
Agents
An agent is a logical component that makes decisions. In Risicare, agents are identified by:
- ID: Unique identifier (auto-generated or explicit)
- Name: Human-readable name (e.g., "planner", "researcher")
- Role: The agent's role, as free text. The
AgentRoleenum gives canonical values: 12 in Python (for exampleorchestrator,worker,critic,validator) and 14 in JavaScript, which addsreviewerandcustom - Type: The agent framework/pattern used
@agent(name="planner", role="orchestrator")
def plan_task(objective):
# All spans inside this function are associated with this agent
passSessions
A session groups related traces from the same user interaction. Use sessions to:
- Track multi-turn conversations
- Group related agent executions
- Analyze user journeys
with session_context(session_id="user-123-session"):
# All traces here belong to this session
result1 = agent.run("First request")
result2 = agent.run("Follow-up request")Semantic Phases
Risicare tracks semantic phases to understand agent decision-making:
| Phase | Description | Example |
|---|---|---|
| THINK | Reasoning and planning | Analyzing the problem |
| DECIDE | Making a decision | Choosing which tool to use |
| ACT | Taking an action | Calling an API |
| OBSERVE | Reading state | Checking memory |
The SemanticPhase enum has three more values: REFLECT, COMMUNICATE and COORDINATE.
@trace_think
def analyze_problem(context):
"""This is a THINK phase - reasoning about the problem"""
pass
@trace_decide
def choose_tool(options):
"""This is a DECIDE phase - selecting an action"""
pass
@trace_act
def execute_tool(tool, args):
"""This is an ACT phase - performing the action"""
passProgressive Integration
Risicare supports incremental adoption. Start with one import and two environment variables, and add richer observability as needed:
| Tier | Code | What You Get |
|---|---|---|
| 0 | import risicare + RISICARE_API_KEY + RISICARE_TRACING=true | LLM calls of the supported providers traced |
| 1 | risicare.init(...) | Explicit configuration |
| 2 | @agent() | Agent identity tracking |
| 3 | @session | User session grouping |
| 4 | @trace_think/decide/act | Decision phase tracking |
| 5 | @trace_message/delegate | Full multi-agent support |
Context Propagation
Risicare automatically propagates context through your code:
- Thread-safe: Uses Python
contextvarsfor thread isolation - Async-safe: Works correctly with
asyncio - Cross-process: Not automatic. Use
inject_trace_context()andextract_trace_context()to carry W3C Trace Context headers between services
Automatic Propagation
@agent(name="parent")
async def parent_agent():
# Context automatically propagates to child calls
await child_agent() # Inherits trace context
@agent(name="child")
async def child_agent():
# This agent's spans are children of parent_agent's span
passManual Context
# Extract context for passing to another system
context = get_trace_context()
# Restore context in another thread/process
with restore_trace_context(context):
# Spans created here continue the trace
passError Taxonomy
When errors occur, Risicare classifies them using a 10-module taxonomy:
| Module | What It Covers |
|---|---|
| PERCEPTION | Input parsing, validation |
| REASONING | Logic errors, hallucinations |
| TOOL | Tool execution failures |
| MEMORY | State management issues |
| OUTPUT | Response formatting |
| COORDINATION | Workflow problems |
| COMMUNICATION | Inter-agent messages |
| ORCHESTRATION | Agent lifecycle |
| CONSENSUS | Multi-agent agreement |
| RESOURCES | Resource contention |
Each module contains categories, and each category contains specific error codes:
TOOL.EXECUTION.TIMEOUT
│ │ └── Specific error code
│ └── Category (EXECUTION)
└── Module (TOOL)
The Diagnosis Pipeline
Risicare's diagnosis pipeline has four stages. It is built, and it is held for the beta — it does not run for new errors:
Error Detected
↓
1. Context Extraction
Extract relevant spans, messages, and state
↓
2. Taxonomy Classification
Classify using Llama-3.3-70B (via Together.AI)
↓
3. Root Cause Analysis
Deep analysis using Llama-3.3-70B (via Together.AI)
↓
4. Fix Suggestion
Generate fix configurations
Held for the beta
None of these four stages runs for a new error today. A diagnosis cannot be requested, and automatic diagnosis is off by default, so no new root cause or fix configuration is produced. There is no way to read earlier diagnoses or fixes with an API key during the beta, and the dashboard does not show them. The error classification described under Error Taxonomy is separate and runs on every failure the SDK records. Deploying fixes via A/B testing is built and gated off, and hypothesis testing is built but not connected to the runtime.