Skip to main content
GitHub

Evaluations

Built-in scorers for LLM outputs — running an evaluation is held for the beta.

Running evaluations is held for the beta

Evaluations never run automatically on ingest, and an evaluation cannot be started in this release. In the dashboard, Evaluations is removed from the navigation, and the page shows only a notice that evaluations are not part of this release. There is no way to list or read evaluations with an API key during the beta. The scorers below are implemented and describe what an evaluation measures.

Risicare includes 13 built-in scorers for evaluating LLM outputs across RAG, safety, agent behavior, and general quality.

Overview

As designed, an evaluation is requested through the API or the dashboard's Evaluations page, naming the scorers to run against a set of traces.

Scorer Categories

RAG Scorers

Evaluate retrieval-augmented generation quality:

ScorerClassDescription
faithfulnessFaithfulnessScorerDoes the response stay faithful to retrieved context?
answer_relevancyAnswerRelevancyScorerIs the response relevant to the query?
context_precisionContextPrecisionScorerHow precise is the retrieved context?
context_recallContextRecallScorerDoes the context contain all needed information?
hallucinationHallucinationScorerDoes the response contain hallucinated information?

Safety Scorers

Detect harmful or inappropriate content:

ScorerClassDescription
toxicityToxicityScorerOffensive or harmful language
biasBiasScorerUnfair or prejudiced content
pii_leakagePIILeakageScorerPersonal identifiable information leakage

Agent Scorers

Evaluate agent behavior:

ScorerClassDescription
tool_correctnessToolCorrectnessScorerDid the agent select appropriate tools?
task_completionTaskCompletionScorerDid the agent complete the task?
goal_accuracyGoalAccuracyScorerHow accurately did the agent achieve the goal?

General Scorers

General quality metrics:

ScorerClassDescription
g_evalGEvalScorerGeneral evaluation (coherence, structure, quality)
factualityFactualityScorerIs the response factually correct?

ScorerInput Fields

All scorers accept a ScorerInput dataclass. Each scorer uses a subset of these fields based on its required_fields.

FieldTypeDescription
trace_idstrUnique identifier for the trace being evaluated (required)
span_idstr | NoneOptional span identifier within the trace
questionstr | NoneThe user's question/query (RAG scorers)
answerstr | NoneThe AI's response/answer to evaluate (RAG scorers)
contextslist[str]List of context passages retrieved for RAG
ground_truthstr | NoneThe expected/correct answer for comparison
expected_toolslist[str]List of tool names the agent should have used
used_toolslist[str]List of tool names the agent actually used
tool_callslist[dict]Detailed tool call information with parameters
task_descriptionstr | NoneDescription of the task assigned to the agent
goalstr | NoneThe goal the agent was trying to achieve
output_textstr | NoneGeneric output text to evaluate
input_textstr | NoneGeneric input text for context
custom_criteriastr | NoneUser-defined evaluation criteria (G-Eval)
evaluation_stepslist[str]Steps to follow during evaluation (G-Eval)
metadatadictAdditional key-value metadata

Required Fields by Scorer

ScorerRequired Fields
faithfulnessanswer, contexts
answer_relevancyquestion, answer
context_precisionquestion, contexts
context_recallcontexts, ground_truth
hallucinationanswer, contexts
toxicityoutput_text
biasoutput_text
pii_leakageoutput_text
tool_correctness(none required, uses expected_tools and used_tools)
task_completiontask_description, output_text
goal_accuracygoal, output_text
g_evaloutput_text
factualityoutput_text

Requesting an Evaluation

API

Not available with an API key during the beta. As designed, an evaluation request names the evaluation, its type, the traces to evaluate and the scorers to run (criteria).

Dashboard

While evaluations are held for the beta, the dashboard's Evaluations page shows only a notice that evaluations are not part of this release, and it is removed from the navigation.

Evaluation Results

Results include:

{
  "trace_id": "abc123",
  "evaluations": [
    {
      "scorer": "faithfulness",
      "score": 0.92,
      "passed": true,
      "reasoning": "Response accurately reflects the retrieved context..."
    },
    {
      "scorer": "toxicity",
      "score": 0.01,
      "passed": true,
      "reasoning": "No toxic content detected."
    }
  ]
}

Running Evaluations

init() does not take an evaluations argument — evaluations are not configured in the SDK. As designed, they are requested through the API (held for the beta, see above), naming the scorers to run in the criteria field. See the Scorers reference for the scorer list.

Next Steps