Overview
How Risicare's diagnosis engine works — built, and held for the beta.
Held for the beta
The diagnosis engine is built and held: it does not run for new errors. A diagnosis cannot be requested, and automatic diagnosis is off by default, with nothing in the product that turns it on. Errors are still detected and classified against the 154-code taxonomy by the SDK as spans are captured. There is no way to read an earlier diagnosis with an API key during the beta. This page describes the engine as designed.
Risicare's diagnosis engine is built to analyze agent failures in a 4-stage LLM-powered pipeline.
How It Works
As designed, when an error occurs in your agent, Risicare:
- Detects the error from trace data
- Extracts relevant context (spans, messages, tool I/O)
- Classifies using the error taxonomy
- Suggests potential fixes
Error Detected → Context Extraction → Classification → Fix Suggestion
(auto) (100ms) (1-2s) (500ms)
Diagnosis Pipeline
Stage 1: Context Extraction
Extract relevant information from the error trace:
- Error span and parent spans
- Recent LLM prompts and completions
- Tool inputs and outputs
- Agent state and messages
- Surrounding context (before/after)
Max context: 50 spans, 100K tokens
Stage 2: Taxonomy Classification
Classify the error using the 10-module taxonomy (154 error codes across 31 categories):
| Module | Focus Area |
|---|---|
| PERCEPTION | Input processing |
| REASONING | Logic and inference |
| TOOL | Tool execution |
| MEMORY | State management |
| OUTPUT | Response generation |
| COORDINATION | Workflow control |
| COMMUNICATION | Inter-agent messages |
| ORCHESTRATION | Lifecycle management |
| CONSENSUS | Agreement protocols |
| RESOURCES | Shared resource access |
Classification uses a heuristic-first approach: a pattern matcher with 379 rules attempts to classify the error before any LLM call. A regex match assigns a fixed confidence of 0.9 and the LLM step is skipped; if no rule matches, the error falls back to the LLM classifier. This keeps typical classification under 100ms and reduces LLM costs.
LLM fallback: meta-llama/Llama-3.3-70B-Instruct-Turbo via Together.AI (used only when no pattern matches). gpt-4o-mini is used instead only when an OpenAI key is configured (production has none).
Stage 3: Root Cause Analysis
Deep analysis of why the error occurred:
- Identify contributing factors
- Trace causal chain
- Distinguish symptoms from causes
- Assess severity and impact
Model: meta-llama/Llama-3.3-70B-Instruct-Turbo via Together.AI (detailed reasoning). When an OpenAI key is configured, gpt-4o is used instead as a fallback; production runs the Together.AI default.
Stage 4: Fix Suggestion
Recommend fixes based on the diagnosis:
- Template-first: deterministic fix templates keyed by
error_codeare applied when one exists - LLM fallback: generate a new fix from the root cause when no template matches
- Rank fixes by confidence (minimum 0.5 to be included)
Stored knowledge base is planned, not built
A persistent cross-diagnosis knowledge base (the knowledge_patterns / fix_templates tables) is planned but not yet implemented — those tables do not exist in production. Stage 4 today is template-first plus LLM fallback only.
Diagnosis Output
A completed diagnosis, as designed, carries the error classification, the root cause analysis and the suggested fixes (abbreviated). You cannot read one with an API key during the beta:
{
"diagnosis_id": "diag-abc123",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "00f067aa0ba902b7",
"status": "completed",
"error_code": "TOOL.EXECUTION.TIMEOUT",
"classification": {
"module": "TOOL",
"category": "EXECUTION",
"subcategory": "TIMEOUT",
"error_code": "TOOL.EXECUTION.TIMEOUT",
"confidence": 0.92
},
"root_cause": {
"root_cause": "External API timeout due to large payload",
"contributing_factors": [
"Payload size: 2.5MB exceeds typical 100KB",
"No timeout configured on API call",
"Single retry with no backoff"
],
"confidence": 0.85
},
"suggested_fixes": [
{
"fix_type": "retry",
"title": "Add exponential backoff retry",
"description": "Retry the call with exponential backoff",
"confidence": 0.85,
"fix_config": {
"max_retries": 3,
"initial_delay_ms": 1000,
"exponential_base": 2.0,
"max_delay_ms": 30000,
"jitter": true,
"retry_on": []
}
},
{
"fix_type": "parameter",
"title": "Increase timeout to 60s",
"description": "Raise the tool call timeout",
"confidence": 0.70,
"fix_config": {
"timeout_ms": 60000
}
}
]
}No evidence field
A root cause carries root_cause, contributing_factors and confidence.
There is no evidence or reasoning_chain field.
Triggering Diagnosis
Both ways of requesting a diagnosis are held for the beta.
Automatic
Automatic diagnosis is off by default, and nothing in the product turns it on. As built, it considers spans recorded with an error.
Reporting caught exceptions
Unhandled exceptions are recorded automatically. For exceptions you catch in a try/except, use report_error() to record them on the trace:
from risicare import report_error
try:
result = tool.execute()
except ToolError as e:
report_error(e) # Records the error on the trace
result = fallback()report_error() never raises. Inside a traced context it records the error on the current span. Outside any trace it creates a standalone error span with automatic deduplication (same error type + message suppressed for 5 minutes). It does not start a diagnosis.
Manual
A diagnosis cannot be requested by API during the beta: the Management API is not available with an API key.
The dashboard's Diagnose button is hidden while diagnosis is held.
Diagnosis Caching
Similar errors use cached diagnoses:
- Cache key:
error_code + stack_trace_hash - Cache TTL: 24 hours
- Cache hit rate: ~60% target (the cache is per-process in-memory and not shared across worker processes)
Performance
| Metric | Target |
|---|---|
| Detection to diagnosis | < 5s P50 |
| Classification accuracy | > 90% |
| Fix suggestion relevance | > 80% |
| Cache hit rate | > 50% |