Skip to main content
GitHub

Diagnose

Error classification, and the LLM-powered diagnosis pipeline held for the beta.

Diagnosis is held for the beta

What runs today: the SDK classifies an exception that leaves a traced call against the 154-code taxonomy. A span that is only marked as an error gets no code (details). The SDK assigns 19 of the codes, the same in Python and JavaScript; an error that matches no rule gets TOOL.EXECUTION.CRASHED.

What does not run yet: the deeper diagnosis pass — root-cause analysis and ranked fix suggestions. A diagnosis cannot be requested, and automatic diagnosis is off by default (auto_diagnosis_enabled). Those are the only two ways a diagnosis gets produced, so no new diagnosis is created today. There is no way to read an earlier diagnosis with an API key during the beta, and the dashboard does not show diagnoses while diagnosis is held. Everything downstream — fix generation, hypothesis testing, deployment — waits on it.

The pages in this section describe the pipeline as designed and are kept so you can see what is coming. Nothing here is deleted; it is not yet reachable.

Risicare's diagnosis pipeline is built to analyze errors in your AI agents in four LLM-powered stages.

Beyond observability

Most AI observability platforms stop at showing you errors. Risicare classifies every failure against a 154-code taxonomy as it is captured, and the diagnosis pipeline adds root-cause analysis on top of that once it is enabled.

Overview

As designed, when an error occurs the diagnosis engine:

  1. Extracts context from the trace (spans, messages, tool I/O)
  2. Classifies the error using the 10-module taxonomy
  3. Analyzes root cause with deep LLM reasoning
  4. Suggests fixes based on patterns and templates

The 4-Stage Pipeline

Error Detected (span with has_error=true)
           ↓
┌─────────────────────────────┐
│ Stage 1: Context Extraction │  Extract relevant spans,
│         (~100ms)            │  messages, and state
└─────────────────────────────┘
           ↓
┌─────────────────────────────┐
│ Stage 2: Classification     │  Classify using Llama-3.3-70B
│         (~500ms)            │  Fast, cheap categorization
└─────────────────────────────┘
           ↓
┌─────────────────────────────┐
│ Stage 3: Root Cause         │  Deep analysis using Llama-3.3-70B
│         (~2-3s)             │  Thorough reasoning
└─────────────────────────────┘
           ↓
┌─────────────────────────────┐
│ Stage 4: Fix Suggestion     │  Generate fix configs
│         (~500ms)            │  Pattern match first
└─────────────────────────────┘
           ↓
    DiagnosisResult

Error Taxonomy

Risicare classifies errors into a hierarchical taxonomy:

ModuleWhat It CoversExample Codes
PERCEPTIONInput parsing, validationPERCEPTION.PARSING.JSON_INVALID
REASONINGLogic errors, hallucinationsREASONING.LOGIC.CONTRADICTION
TOOLTool execution failuresTOOL.EXECUTION.TIMEOUT
MEMORYState management issuesMEMORY.RETRIEVAL.NOT_FOUND
OUTPUTResponse formattingOUTPUT.FORMAT.INVALID_JSON
COORDINATIONWorkflow problemsCOORDINATION.WORKFLOW.DEADLOCK
COMMUNICATIONInter-agent messagesCOMMUNICATION.ROUTING.NO_ROUTE
ORCHESTRATIONAgent lifecycleORCHESTRATION.LIFECYCLE.HUNG
CONSENSUSMulti-agent agreementCONSENSUS.AGREEMENT.QUORUM_FAILED
RESOURCESResource contentionRESOURCES.ACCESS.LOCKED

Error Code Format

Error codes follow the pattern: MODULE.CATEGORY.SPECIFIC_ERROR

Example: TOOL.EXECUTION.TIMEOUT means:

  • Module: TOOL (action errors)
  • Category: EXECUTION (runtime execution)
  • Specific: TIMEOUT (operation timed out)

Diagnosis Result

A diagnosis, as designed (abbreviated). You cannot read one with an API key during the beta:

{
  "diagnosis_id": "diag-abc123",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "status": "completed",
  "error_code": "TOOL.EXECUTION.TIMEOUT",
  "classification": {
    "module": "TOOL",
    "category": "EXECUTION",
    "subcategory": "TIMEOUT",
    "error_code": "TOOL.EXECUTION.TIMEOUT",
    "confidence": 0.92
  },
  "root_cause": {
    "root_cause": "The weather API call timed out after 30 seconds due to network latency",
    "contributing_factors": [],
    "confidence": 0.85
  },
  "suggested_fixes": [
    {
      "fix_type": "retry",
      "title": "Retry with exponential backoff",
      "description": "Retry the call up to 3 times with exponential backoff",
      "confidence": 0.85,
      "fix_config": {"max_retries": 3, "initial_delay_ms": 1000, "exponential_base": 2.0}
    },
    {
      "fix_type": "parameter",
      "title": "Increase timeout",
      "description": "Raise the timeout to 60 seconds",
      "confidence": 0.72,
      "fix_config": {"timeout_ms": 60000}
    }
  ]
}

Automatic vs Manual Diagnosis

Both ways of requesting a diagnosis are held for the beta.

Automatic

Automatic diagnosis is off by default, and nothing in the product turns it on. As built, it considers spans recorded with an error (has_error=true).

Manual

A diagnosis cannot be requested by API during the beta: the Management API is not available with an API key.

The dashboard's Diagnose button is hidden while diagnosis is held.

Caching

Similar errors are cached for 24 hours:

  • Same error_code + similar stack trace = cache hit
  • Avoids redundant LLM calls
  • Cache can be invalidated manually

The cache is per-process in-memory (not shared across worker processes), so hit rates are a target rather than a guarantee.

Next Steps