Skip to main content
GitHub

Heal

Fix generation and deployment for AI agent failures — built, and held for the beta.

Fix generation is held for the beta

Fix generation runs downstream of a diagnosis, and diagnosis is currently held: a diagnosis cannot be requested and automatic diagnosis is off by default. With no new diagnosis, no new fix is generated today. There is no way to read earlier fixes with an API key during the beta, and the dashboard does not show them.

Fix deployment, A/B testing and rollback are built but gated off as well — a fix cannot be promoted and a deployment cannot be created, and that machinery has never run on real traffic. The knowledge base exists in the codebase but is not connected to anything.

These pages describe the pipeline as designed and are kept so you can see what is coming. Nothing here is deleted; it is not yet reachable.

Risicare detects and classifies errors automatically today. Diagnosing root causes and generating fix recommendations are built and held — see the callout above.

Beyond observability

Risicare is built to go past showing you the error: a 154-code taxonomy classifies every failure today, and the diagnosis and fix-generation stages described below explain why it happened and propose a fix once they are enabled.

Overview

The healing pipeline follows the DoVer methodology (Diagnosis via Observation of Verification):

  1. Generate Hypotheses - Create testable hypotheses about fixes
  2. Validate Statistically (built, never exercised) - Test fixes with A/B testing
  3. Deploy Safely (built, never exercised) - Canary release with automatic rollback

Fix Types

Risicare can generate 7 types of fixes:

TypeWhat It DoesExample
PromptModify system prompt or add few-shot examplesAdd clarifying instructions
ParameterAdjust LLM parametersLower temperature, increase max_tokens
ToolFix tool configurationAdd timeout, fix validation
RetryAdd retry logicExponential backoff on transient errors
FallbackUse alternative model/strategyFall back to gpt-4o-mini on timeout
GuardAdd input/output validationJSON schema validation
RoutingChange agent delegationRoute to different specialist agent

Fix Configuration

Fixes are JSON configurations, not code:

{
  "fix_id": "fix-abc123",
  "fix_type": "retry",
  "config": {
    "max_retries": 3,
    "initial_delay_ms": 1000,
    "exponential_base": 2.0,
    "max_delay_ms": 30000,
    "jitter": true,
    "retry_on": ["TimeoutError"]
  }
}

No Code Injection

Fixes are declarative configurations applied by the SDK at runtime. Risicare never injects code into your system.

Hypothesis Testing

Before deployment, fixes are validated through hypothesis testing:

Generate Hypotheses

Diagnosis: TOOL.EXECUTION.TIMEOUT on weather_api

Hypothesis 1: Adding retry with backoff will reduce timeout errors
  Prior probability: 0.75 (based on similar patterns)

Hypothesis 2: Increasing timeout to 60s will reduce errors
  Prior probability: 0.60

Hypothesis 3: Adding fallback to cached data will maintain uptime
  Prior probability: 0.55

Statistical Validation

Each hypothesis is tested with:

  • Sample size calculation for statistical power (0.8)
  • Two-proportion z-test for significance (p < 0.05)
  • Bayesian updates to posterior probability
  • O'Brien-Fleming boundaries for early stopping
Test Results:
  Baseline error rate: 12.3%
  Treatment error rate: 2.1%
  Effect size (Cohen's h): 0.38
  P-value: 0.0023 ✓

  Decision: Hypothesis VALIDATED

Deployment Pipeline

Fix Created
     ↓
┌─────────────────┐
│ Canary (5%)     │  Minimum 100 samples
│                 │  Monitor error rate
└─────────────────┘
     ↓ (if passing)
┌─────────────────┐
│ Ramp (25%)      │  Statistical A/B test
│                 │  O'Brien-Fleming boundaries
└─────────────────┘
     ↓ (if winning)
┌─────────────────┐
│ Ramp (50%)      │  Continue testing
│                 │
└─────────────────┘
     ↓ (if winning)
┌─────────────────┐
│ Graduate (100%) │  Hold for 24 hours
│                 │  Mark as graduated
└─────────────────┘

Automatic Rollback

Fixes are automatically rolled back if:

  • Error rate increases >10% vs baseline
  • P99 latency exceeds 2x baseline
  • Manual rollback triggered

Rollback latency design target: under 500ms — never measured. Automatic rollback runs in the deployment worker, which is gated off with the rest of fix deployment, so it does not run today.

Fix Runtime

The SDK includes a fix runtime. It is off by default since Python 0.4.0 and JavaScript 0.7.0. As designed, when it is on it:

  1. Loads fixes from the API on startup
  2. Caches locally with periodic refresh
  3. Routes requests based on A/B assignment
  4. Applies fixes at LLM call time
import risicare
 
# init() does not start the fix runtime: fix_runtime defaults to False since 0.4.0.
risicare.init(api_key="rsk-...")
 
# To start it, pass fix_runtime=True (or set RISICARE_FIX_RUNTIME=true).
# During the beta the route that the runtime reads is not reachable with an
# API key, so the runtime loads no fix. Keep it off.

Knowledge Base (planned)

A knowledge base (stage 6) is designed to store successful fixes so similar errors can reuse a known remedy instead of regenerating one:

  • Error patterns and fix templates keyed by error code
  • Cross-customer learning (federated, no raw data)

This stage is not connected — a knowledge store exists in the codebase but nothing uses it, there is no fix-template or pattern store in production, and fix generation (held for the beta) uses templates and LLM fallback only (see Diagnose → Pipeline).

Next Steps