Diagnose
Error classification, and the LLM-powered diagnosis pipeline held for the beta.
Diagnosis is held for the beta
What runs today: the SDK classifies an exception that leaves a traced call
against the 154-code taxonomy. A span that is only marked as an error gets no
code (details). The SDK assigns 19 of the codes, the same in
Python and JavaScript; an error that matches no rule gets
TOOL.EXECUTION.CRASHED.
What does not run yet: the deeper diagnosis pass — root-cause analysis and
ranked fix suggestions. A diagnosis cannot be requested, and automatic diagnosis
is off by default (auto_diagnosis_enabled). Those are the only two ways a
diagnosis gets produced, so no new diagnosis is created today. There is no way
to read an earlier diagnosis with an API key during the beta, and the dashboard
does not show diagnoses while diagnosis is held. Everything downstream — fix
generation, hypothesis testing, deployment — waits on it.
The pages in this section describe the pipeline as designed and are kept so you can see what is coming. Nothing here is deleted; it is not yet reachable.
Risicare's diagnosis pipeline is built to analyze errors in your AI agents in four LLM-powered stages.
Beyond observability
Most AI observability platforms stop at showing you errors. Risicare classifies every failure against a 154-code taxonomy as it is captured, and the diagnosis pipeline adds root-cause analysis on top of that once it is enabled.
Overview
As designed, when an error occurs the diagnosis engine:
- Extracts context from the trace (spans, messages, tool I/O)
- Classifies the error using the 10-module taxonomy
- Analyzes root cause with deep LLM reasoning
- Suggests fixes based on patterns and templates
Error Taxonomy
10 modules, 31 categories, 154 error codes
Diagnosis Pipeline
How the 4-stage pipeline works
Overview
How diagnosis works
The 4-Stage Pipeline
Error Detected (span with has_error=true)
↓
┌─────────────────────────────┐
│ Stage 1: Context Extraction │ Extract relevant spans,
│ (~100ms) │ messages, and state
└─────────────────────────────┘
↓
┌─────────────────────────────┐
│ Stage 2: Classification │ Classify using Llama-3.3-70B
│ (~500ms) │ Fast, cheap categorization
└─────────────────────────────┘
↓
┌─────────────────────────────┐
│ Stage 3: Root Cause │ Deep analysis using Llama-3.3-70B
│ (~2-3s) │ Thorough reasoning
└─────────────────────────────┘
↓
┌─────────────────────────────┐
│ Stage 4: Fix Suggestion │ Generate fix configs
│ (~500ms) │ Pattern match first
└─────────────────────────────┘
↓
DiagnosisResult
Error Taxonomy
Risicare classifies errors into a hierarchical taxonomy:
| Module | What It Covers | Example Codes |
|---|---|---|
| PERCEPTION | Input parsing, validation | PERCEPTION.PARSING.JSON_INVALID |
| REASONING | Logic errors, hallucinations | REASONING.LOGIC.CONTRADICTION |
| TOOL | Tool execution failures | TOOL.EXECUTION.TIMEOUT |
| MEMORY | State management issues | MEMORY.RETRIEVAL.NOT_FOUND |
| OUTPUT | Response formatting | OUTPUT.FORMAT.INVALID_JSON |
| COORDINATION | Workflow problems | COORDINATION.WORKFLOW.DEADLOCK |
| COMMUNICATION | Inter-agent messages | COMMUNICATION.ROUTING.NO_ROUTE |
| ORCHESTRATION | Agent lifecycle | ORCHESTRATION.LIFECYCLE.HUNG |
| CONSENSUS | Multi-agent agreement | CONSENSUS.AGREEMENT.QUORUM_FAILED |
| RESOURCES | Resource contention | RESOURCES.ACCESS.LOCKED |
Error Code Format
Error codes follow the pattern: MODULE.CATEGORY.SPECIFIC_ERROR
Example: TOOL.EXECUTION.TIMEOUT means:
- Module: TOOL (action errors)
- Category: EXECUTION (runtime execution)
- Specific: TIMEOUT (operation timed out)
Diagnosis Result
A diagnosis, as designed (abbreviated). You cannot read one with an API key during the beta:
{
"diagnosis_id": "diag-abc123",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "00f067aa0ba902b7",
"status": "completed",
"error_code": "TOOL.EXECUTION.TIMEOUT",
"classification": {
"module": "TOOL",
"category": "EXECUTION",
"subcategory": "TIMEOUT",
"error_code": "TOOL.EXECUTION.TIMEOUT",
"confidence": 0.92
},
"root_cause": {
"root_cause": "The weather API call timed out after 30 seconds due to network latency",
"contributing_factors": [],
"confidence": 0.85
},
"suggested_fixes": [
{
"fix_type": "retry",
"title": "Retry with exponential backoff",
"description": "Retry the call up to 3 times with exponential backoff",
"confidence": 0.85,
"fix_config": {"max_retries": 3, "initial_delay_ms": 1000, "exponential_base": 2.0}
},
{
"fix_type": "parameter",
"title": "Increase timeout",
"description": "Raise the timeout to 60 seconds",
"confidence": 0.72,
"fix_config": {"timeout_ms": 60000}
}
]
}Automatic vs Manual Diagnosis
Both ways of requesting a diagnosis are held for the beta.
Automatic
Automatic diagnosis is off by default, and nothing in the product turns it on. As built, it considers spans recorded with an error (has_error=true).
Manual
A diagnosis cannot be requested by API during the beta: the Management API is not available with an API key.
The dashboard's Diagnose button is hidden while diagnosis is held.
Caching
Similar errors are cached for 24 hours:
- Same
error_code+ similar stack trace = cache hit - Avoids redundant LLM calls
- Cache can be invalidated manually
The cache is per-process in-memory (not shared across worker processes), so hit rates are a target rather than a guarantee.