Skip to main content
GitHub

Fix Runtime

Apply self-healing fixes at runtime.

Held for the beta — off by default, and no fix reaches it

The Fix Runtime is built, and it is held for the beta. It is off by default since Python 0.4.0 and JavaScript 0.7.0: init() starts it only with fix_runtime=True (or RISICARE_FIX_RUNTIME=true). The route that it reads, GET /v1/fixes/active on the API host, is not reachable with an API key during the beta, so a runtime that you start logs a warning and loads no fix. Diagnosis is held, so no new fix is generated, and fix deployment is gated off, so no fix can be put into a live status. Keep the Fix Runtime off. This page describes the runtime as built.

The Fix Runtime enables dynamic application of fixes without redeployment.

Overview

Fix Runtime flow: Your Code calls LLM, Fix Runtime checks for active fix, applies modification or passes through, records A/B metrics

When Risicare detects an error pattern and generates a fix, the Fix Runtime can apply it immediately:

from risicare.runtime import FixRuntime, FixRuntimeConfig
 
runtime = FixRuntime(
    config=FixRuntimeConfig(
        api_endpoint="https://api.risicare.ai",
        api_key="rsk-...",
        enabled=True,
    )
)
 
runtime.start()

Getting Started

A complete setup that initializes the fix runtime alongside the SDK:

import risicare
from risicare import init_runtime, get_runtime, shutdown_runtime
from risicare.runtime import FixRuntimeConfig
 
# 1. Initialize SDK (tracing + auto-instrumentation).
#    fix_runtime stays False, so init() starts no runtime.
risicare.init(api_key="rsk-...", environment="production")
 
# 2. Configure fix runtime
config = FixRuntimeConfig(
    api_endpoint="https://api.risicare.ai",
    api_key="rsk-...",
    cache_ttl_seconds=300,        # Cache fixes locally for 5 minutes
    refresh_interval_seconds=60,  # Poll API for new fixes every minute
    dry_run=False,                # Set True to log fixes without applying
    ab_testing_enabled=True,      # Enable A/B traffic splitting
)
 
# 3. Initialize and auto-start (fetches active fixes from API)
runtime = init_runtime(config=config, auto_start=True)
 
# 4. Use anywhere in your application
runtime = get_runtime()
fix = runtime.get_fix("TOOL.EXECUTION.TIMEOUT", session_id="sess-123")
 
# 5. Shutdown on exit (stops background refresh, clears cache)
shutdown_runtime()

On init_runtime(auto_start=True):

  • The FixLoader requests active fixes from /v1/fixes/active on api_endpoint. During the beta this route is not reachable with an API key: the loader logs a warning and loads no fix
  • The FixCache stores them locally (LRU, keyed by error code, TTL-based expiration)
  • A background thread polls the API every refresh_interval_seconds for updates
  • The FixInterceptor chain applies matching fixes when wrap_call() or intercept_call() is used

FixRuntimeConfig

FixRuntimeConfig is a dataclass. Every field has a default, so pass only what you want to override:

from risicare import FixRuntimeConfig
 
config = FixRuntimeConfig(
    api_endpoint="https://api.risicare.ai",
    api_key="rsk-...",
    cache_ttl_seconds=300,
    dry_run=True,
)
FieldTypeDefaultDescription
api_endpointstr""Risicare API endpoint
api_keystr | NoneNoneAPI key
project_idstr | NoneNoneProject ID
cache_enabledboolTrueEnable local fix caching
cache_ttl_secondsint300Cache TTL (5 minutes)
cache_max_entriesint1000Max cached fixes
auto_refreshboolTrueAuto-refresh fixes from API
refresh_interval_secondsint60Refresh interval
enabledboolTrueEnable fix runtime
dry_runboolFalseLog fixes without applying
max_retriesint3Max retries per fix
ab_testing_enabledboolTrueEnable A/B testing
track_effectivenessboolTrueTrack fix effectiveness
timeout_msint1000Fix application timeout
debugboolFalseDebug logging

FixRuntimeConfig is importable from the package root (from risicare import FixRuntimeConfig) or from risicare.runtime.

You can also create config from environment variables:

config = FixRuntimeConfig.from_env()

This reads RISICARE_API_KEY, RISICARE_PROJECT_ID and other RISICARE_* environment variables. For api_endpoint it reads RISICARE_API_ENDPOINT, then RISICARE_ENDPOINT, and then uses https://api.risicare.ai.

Fix Types

Risicare's fix generator produces 7 fix types, but the runtime can only apply 5 of them. tool and routing fixes have no applier in the SDK:

TypeDescriptionApplied by the runtime?
promptSystem prompt modificationsYes — on pre_call
parameterLLM parameter adjustments (temperature, etc.)Yes — on pre_call
retryRetry logic with backoffYes — on on_error
fallbackAlternative model/provider fallbackYes — on pre_call / on_error
guardInput/output validation guardsYes — on pre_call / post_call
toolTool configuration fixesNo — not implemented
routingRequest routing changesNo — not implemented

tool and routing fixes are silently ignored

The interceptor chain has no branch for tool or routing. If a fix of either type reaches the runtime it is dropped without applying, without raising, and without recording an entry in applied_fixes — so nothing in the SDK signals that the fix did not take effect. Treat generated tool and routing fixes as recommendations to apply by hand.

Public Methods

start() / stop()

Start and stop the runtime lifecycle:

runtime = FixRuntime(config=config)
 
# Start: loads fixes from API and begins background refresh
runtime.start()
 
# ... your application runs ...
 
# Stop: halts background refresh and clears cache
runtime.stop()

get_fix()

Look up the active fix for a given error code:

fix = runtime.get_fix(
    error_code="TOOL.EXECUTION.TIMEOUT",
    session_id="session-123",  # optional, used for A/B bucketing
)
 
if fix:
    print(f"Fix {fix.fix_id} (type: {fix.fix_type})")

wrap_call()

Wrap a synchronous operation with fix interception and retry handling. This is a convenience method that returns a new callable -- it is not a decorator.

def call_llm(messages: list) -> str:
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=messages
    )
    return response.choices[0].message.content
 
wrapped_fn = runtime.wrap_call(
    call_llm,
    operation_type="llm_call",
    operation_name="gpt-4o",
    session_id="session-123",
    max_retries=3,
)
 
# Call the wrapped function -- during the beta the runtime loads no fix (the
# route that it reads is not reachable with an API key), so the call runs
# unchanged.
result = wrapped_fn(messages=[{"role": "user", "content": "Hello"}])

wrap_call is not a decorator

wrap_call returns a wrapped callable. Do not use it as @runtime.wrap_call. Pass the function as the first argument and call the returned wrapper.

wrap_async_call()

Async version of wrap_call:

async def call_llm_async(messages: list) -> str:
    response = await async_client.chat.completions.create(
        model="gpt-4o",
        messages=messages
    )
    return response.choices[0].message.content
 
wrapped_fn = runtime.wrap_async_call(  # sync — returns an async callable
    call_llm_async,
    operation_type="llm_call",
    operation_name="gpt-4o",
    session_id="session-123",
    max_retries=3,
)
 
result = await wrapped_fn(messages=[{"role": "user", "content": "Hello"}])

intercept_call() / intercept_response() / intercept_error()

Low-level interception methods for manual control:

# Pre-call: apply prompt and parameter fixes
messages, params, context = runtime.intercept_call(
    operation_type="llm_call",
    operation_name="gpt-4o",
    messages=[{"role": "user", "content": "Hello"}],
    params={"temperature": 0.7},
    session_id="session-123",
    error_code="TOOL.EXECUTION.TIMEOUT",  # if retrying after error
)
 
# Make the call with modified messages/params
try:
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=messages,
        **params,
    )
 
    # Post-call: apply output guards
    response, should_continue = runtime.intercept_response(context, response)
 
except Exception as e:
    # Error: decide whether to retry
    should_retry, modified_params = runtime.intercept_error(context, e)
    if should_retry:
        # Retry with modified_params
        pass

refresh_fixes() / refresh_fixes_async()

Manually refresh fixes from the API (sync and async variants):

# Synchronous
fixes = runtime.refresh_fixes()
 
# Asynchronous
fixes = await runtime.refresh_fixes_async()

get_effectiveness_stats()

Get fix success rates tracked by the runtime:

stats = runtime.get_effectiveness_stats()
for fix_id, data in stats.items():
    print(f"{fix_id}: {data['success_rate']:.1%} "
          f"({data['successes']}/{data['applications']})")

add_interceptor()

Add a custom interceptor to the chain:

runtime.add_interceptor(MyCustomInterceptor())

Custom Interceptors

Create custom interceptors by subclassing FixInterceptor:

from risicare.runtime.interceptors import FixInterceptor, InterceptContext
 
class MyInterceptor(FixInterceptor):
    def pre_call(self, context: InterceptContext, messages=None, params=None):
        """Called before an LLM/tool call. Return (messages, params)."""
        if context.operation_name == "gpt-4o":
            params = params or {}
            params["temperature"] = 0.3
        return messages, params
 
    def post_call(self, context: InterceptContext, response):
        """Called after a call. Return (response, should_continue)."""
        # Validate output, transform response, etc.
        return response, True
 
    def on_error(self, context: InterceptContext, error: Exception):
        """Called on error. Return (should_retry, modified_params)."""
        if "timeout" in str(error).lower():
            return True, {"timeout": 60000}
        return False, None
 
runtime.add_interceptor(MyInterceptor())

InterceptContext

The context object passed through all interceptor methods:

FieldTypeDescription
operation_typestr"llm_call", "tool_call", "agent_delegate"
operation_namestrModel name, tool name, or agent name
session_idstr | NoneSession ID for A/B bucketing
trace_idstr | NoneCurrent trace ID
span_idstr | NoneCurrent span ID
error_codestr | NoneError code (when retrying)
error_messagestr | NoneError message (when retrying)
attemptintCurrent attempt number (starts at 1)
applied_fixeslist[ApplyResult]Fixes applied during this intercept chain

A/B Testing

A/B testing is controlled by two settings:

  1. ab_testing_enabled on FixRuntimeConfig -- enables the A/B testing system
  2. traffic_percentage on each ActiveFix -- controls what percentage of sessions receive the fix

Traffic splitting uses session-hash-based bucketing, ensuring the same session always gets the same treatment:

config = FixRuntimeConfig(
    ab_testing_enabled=True,  # Enable A/B testing
)

The runtime checks each fix's traffic_percentage and uses hashlib.md5(session_id) for deterministic bucketing (consistent across process restarts). As designed, results appear in the dashboard with:

  • Error rate comparison (control vs treatment)
  • Latency impact
  • Statistical confidence

No dashboard page shows these results during the beta.

Rollback

As designed, rollback is managed through the deployment API, not through the SDK runtime. The deployment API is not available with an API key during the beta.

As designed, automatic rollback triggers when:

ConditionThreshold
Error rate increase>10% vs baseline
P99 latency increase>2x baseline

The SDK picks up a rollback only on its next refresh cycle (default: 60 seconds); there is no push path.

Fix Lifecycle

The lifecycle as designed (held for the beta):

1. Error detected -> Diagnosis runs
2. Fix generated -> Appears in dashboard
3. Fix approved -> Synced to runtime via API
4. A/B test -> traffic_percentage controls split
5. Statistical validation
6. Full rollout or rollback

Global Runtime

For convenience, you can use the global runtime functions:

from risicare import init_runtime, get_runtime, shutdown_runtime
 
# Initialize and start
runtime = init_runtime(config=config, auto_start=True)
 
# Get the global instance anywhere
runtime = get_runtime()
 
# Shutdown
shutdown_runtime()

Next Steps