Skip to main content
GitHub

Core Concepts

Understand the key concepts in Risicare observability.

This guide explains the core concepts you'll encounter when using Risicare.

Traces and Spans

Traces

A trace represents a complete execution flow through your agent. It starts when your agent receives input and ends when it produces output.

Trace: "Answer user question about weather"
├── Span: parse_input (2ms)
├── Span: llm_call to gpt-4o (1.2s)
├── Span: tool_call to weather_api (300ms)
└── Span: format_response (5ms)

Spans

A span represents a single unit of work within a trace. Spans have:

  • Name: What operation this span represents
  • Kind: The type of span (LLM_CALL, TOOL_CALL, AGENT, etc.)
  • Timing: Start time and duration
  • Attributes: Key-value metadata
  • Events: Named events that occurred during the span
  • Status: UNSET (the default), OK or ERROR

Span Hierarchy

Spans form a tree structure with parent-child relationships:

Create nested spans with the tracer returned by get_tracer(). Each start_span() is a context manager, and a span opened inside another is automatically parented to it:

from risicare import get_tracer, SpanKind
 
tracer = get_tracer()
 
with tracer.start_span("process_request") as parent:
    with tracer.start_span("call_llm", kind=SpanKind.LLM_CALL) as child:
        # child.parent_span_id == parent.span_id
        # child.trace_id      == parent.trace_id
        pass

There is no top-level start_span()

start_span is a method on the tracer, not a module-level export — from risicare import start_span raises ImportError. Always go through get_tracer(). Spans are only recorded after risicare.init(); before that the tracer returns inert no-op spans.

Agents

An agent is a logical component that makes decisions. In Risicare, agents are identified by:

  • ID: Unique identifier (auto-generated or explicit)
  • Name: Human-readable name (e.g., "planner", "researcher")
  • Role: The agent's role, as free text. The AgentRole enum gives canonical values: 12 in Python (for example orchestrator, worker, critic, validator) and 14 in JavaScript, which adds reviewer and custom
  • Type: The agent framework/pattern used
@agent(name="planner", role="orchestrator")
def plan_task(objective):
    # All spans inside this function are associated with this agent
    pass

Sessions

A session groups related traces from the same user interaction. Use sessions to:

  • Track multi-turn conversations
  • Group related agent executions
  • Analyze user journeys
with session_context(session_id="user-123-session"):
    # All traces here belong to this session
    result1 = agent.run("First request")
    result2 = agent.run("Follow-up request")

Semantic Phases

Risicare tracks semantic phases to understand agent decision-making:

PhaseDescriptionExample
THINKReasoning and planningAnalyzing the problem
DECIDEMaking a decisionChoosing which tool to use
ACTTaking an actionCalling an API
OBSERVEReading stateChecking memory

The SemanticPhase enum has three more values: REFLECT, COMMUNICATE and COORDINATE.

@trace_think
def analyze_problem(context):
    """This is a THINK phase - reasoning about the problem"""
    pass
 
@trace_decide
def choose_tool(options):
    """This is a DECIDE phase - selecting an action"""
    pass
 
@trace_act
def execute_tool(tool, args):
    """This is an ACT phase - performing the action"""
    pass

Progressive Integration

Risicare supports incremental adoption. Start with one import and two environment variables, and add richer observability as needed:

TierCodeWhat You Get
0import risicare + RISICARE_API_KEY + RISICARE_TRACING=trueLLM calls of the supported providers traced
1risicare.init(...)Explicit configuration
2@agent()Agent identity tracking
3@sessionUser session grouping
4@trace_think/decide/actDecision phase tracking
5@trace_message/delegateFull multi-agent support

Context Propagation

Risicare automatically propagates context through your code:

  • Thread-safe: Uses Python contextvars for thread isolation
  • Async-safe: Works correctly with asyncio
  • Cross-process: Not automatic. Use inject_trace_context() and extract_trace_context() to carry W3C Trace Context headers between services

Automatic Propagation

@agent(name="parent")
async def parent_agent():
    # Context automatically propagates to child calls
    await child_agent()  # Inherits trace context
 
@agent(name="child")
async def child_agent():
    # This agent's spans are children of parent_agent's span
    pass

Manual Context

# Extract context for passing to another system
context = get_trace_context()
 
# Restore context in another thread/process
with restore_trace_context(context):
    # Spans created here continue the trace
    pass

Error Taxonomy

When errors occur, Risicare classifies them using a 10-module taxonomy:

ModuleWhat It Covers
PERCEPTIONInput parsing, validation
REASONINGLogic errors, hallucinations
TOOLTool execution failures
MEMORYState management issues
OUTPUTResponse formatting
COORDINATIONWorkflow problems
COMMUNICATIONInter-agent messages
ORCHESTRATIONAgent lifecycle
CONSENSUSMulti-agent agreement
RESOURCESResource contention

Each module contains categories, and each category contains specific error codes:

TOOL.EXECUTION.TIMEOUT
 │      │        └── Specific error code
 │      └── Category (EXECUTION)
 └── Module (TOOL)

The Diagnosis Pipeline

Risicare's diagnosis pipeline has four stages. It is built, and it is held for the beta — it does not run for new errors:

Error Detected
     ↓
1. Context Extraction
   Extract relevant spans, messages, and state
     ↓
2. Taxonomy Classification
   Classify using Llama-3.3-70B (via Together.AI)
     ↓
3. Root Cause Analysis
   Deep analysis using Llama-3.3-70B (via Together.AI)
     ↓
4. Fix Suggestion
   Generate fix configurations

Held for the beta

None of these four stages runs for a new error today. A diagnosis cannot be requested, and automatic diagnosis is off by default, so no new root cause or fix configuration is produced. There is no way to read earlier diagnoses or fixes with an API key during the beta, and the dashboard does not show them. The error classification described under Error Taxonomy is separate and runs on every failure the SDK records. Deploying fixes via A/B testing is built and gated off, and hypothesis testing is built but not connected to the runtime.

Next Steps