Skip to main content
GitHub

Overview

How Risicare's diagnosis engine works — built, and held for the beta.

Held for the beta

The diagnosis engine is built and held: it does not run for new errors. A diagnosis cannot be requested, and automatic diagnosis is off by default, with nothing in the product that turns it on. Errors are still detected and classified against the 154-code taxonomy by the SDK as spans are captured. There is no way to read an earlier diagnosis with an API key during the beta. This page describes the engine as designed.

Risicare's diagnosis engine is built to analyze agent failures in a 4-stage LLM-powered pipeline.

How It Works

As designed, when an error occurs in your agent, Risicare:

  1. Detects the error from trace data
  2. Extracts relevant context (spans, messages, tool I/O)
  3. Classifies using the error taxonomy
  4. Suggests potential fixes
Error Detected → Context Extraction → Classification → Fix Suggestion
    (auto)           (100ms)            (1-2s)          (500ms)

Diagnosis Pipeline

Stage 1: Context Extraction

Extract relevant information from the error trace:

  • Error span and parent spans
  • Recent LLM prompts and completions
  • Tool inputs and outputs
  • Agent state and messages
  • Surrounding context (before/after)

Max context: 50 spans, 100K tokens

Stage 2: Taxonomy Classification

Classify the error using the 10-module taxonomy (154 error codes across 31 categories):

ModuleFocus Area
PERCEPTIONInput processing
REASONINGLogic and inference
TOOLTool execution
MEMORYState management
OUTPUTResponse generation
COORDINATIONWorkflow control
COMMUNICATIONInter-agent messages
ORCHESTRATIONLifecycle management
CONSENSUSAgreement protocols
RESOURCESShared resource access

Classification uses a heuristic-first approach: a pattern matcher with 379 rules attempts to classify the error before any LLM call. A regex match assigns a fixed confidence of 0.9 and the LLM step is skipped; if no rule matches, the error falls back to the LLM classifier. This keeps typical classification under 100ms and reduces LLM costs.

LLM fallback: meta-llama/Llama-3.3-70B-Instruct-Turbo via Together.AI (used only when no pattern matches). gpt-4o-mini is used instead only when an OpenAI key is configured (production has none).

Stage 3: Root Cause Analysis

Deep analysis of why the error occurred:

  • Identify contributing factors
  • Trace causal chain
  • Distinguish symptoms from causes
  • Assess severity and impact

Model: meta-llama/Llama-3.3-70B-Instruct-Turbo via Together.AI (detailed reasoning). When an OpenAI key is configured, gpt-4o is used instead as a fallback; production runs the Together.AI default.

Stage 4: Fix Suggestion

Recommend fixes based on the diagnosis:

  • Template-first: deterministic fix templates keyed by error_code are applied when one exists
  • LLM fallback: generate a new fix from the root cause when no template matches
  • Rank fixes by confidence (minimum 0.5 to be included)

Stored knowledge base is planned, not built

A persistent cross-diagnosis knowledge base (the knowledge_patterns / fix_templates tables) is planned but not yet implemented — those tables do not exist in production. Stage 4 today is template-first plus LLM fallback only.

Diagnosis Output

A completed diagnosis, as designed, carries the error classification, the root cause analysis and the suggested fixes (abbreviated). You cannot read one with an API key during the beta:

{
  "diagnosis_id": "diag-abc123",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "status": "completed",
  "error_code": "TOOL.EXECUTION.TIMEOUT",
  "classification": {
    "module": "TOOL",
    "category": "EXECUTION",
    "subcategory": "TIMEOUT",
    "error_code": "TOOL.EXECUTION.TIMEOUT",
    "confidence": 0.92
  },
  "root_cause": {
    "root_cause": "External API timeout due to large payload",
    "contributing_factors": [
      "Payload size: 2.5MB exceeds typical 100KB",
      "No timeout configured on API call",
      "Single retry with no backoff"
    ],
    "confidence": 0.85
  },
  "suggested_fixes": [
    {
      "fix_type": "retry",
      "title": "Add exponential backoff retry",
      "description": "Retry the call with exponential backoff",
      "confidence": 0.85,
      "fix_config": {
        "max_retries": 3,
        "initial_delay_ms": 1000,
        "exponential_base": 2.0,
        "max_delay_ms": 30000,
        "jitter": true,
        "retry_on": []
      }
    },
    {
      "fix_type": "parameter",
      "title": "Increase timeout to 60s",
      "description": "Raise the tool call timeout",
      "confidence": 0.70,
      "fix_config": {
        "timeout_ms": 60000
      }
    }
  ]
}

No evidence field

A root cause carries root_cause, contributing_factors and confidence. There is no evidence or reasoning_chain field.

Triggering Diagnosis

Both ways of requesting a diagnosis are held for the beta.

Automatic

Automatic diagnosis is off by default, and nothing in the product turns it on. As built, it considers spans recorded with an error.

Reporting caught exceptions

Unhandled exceptions are recorded automatically. For exceptions you catch in a try/except, use report_error() to record them on the trace:

from risicare import report_error
 
try:
    result = tool.execute()
except ToolError as e:
    report_error(e)          # Records the error on the trace
    result = fallback()

report_error() never raises. Inside a traced context it records the error on the current span. Outside any trace it creates a standalone error span with automatic deduplication (same error type + message suppressed for 5 minutes). It does not start a diagnosis.

Manual

A diagnosis cannot be requested by API during the beta: the Management API is not available with an API key.

The dashboard's Diagnose button is hidden while diagnosis is held.

Diagnosis Caching

Similar errors use cached diagnoses:

  • Cache key: error_code + stack_trace_hash
  • Cache TTL: 24 hours
  • Cache hit rate: ~60% target (the cache is per-process in-memory and not shared across worker processes)

Performance

MetricTarget
Detection to diagnosis< 5s P50
Classification accuracy> 90%
Fix suggestion relevance> 80%
Cache hit rate> 50%

Next Steps