Evaluations
Built-in scorers for LLM outputs — running an evaluation is held for the beta.
Running evaluations is held for the beta
Evaluations never run automatically on ingest, and an evaluation cannot be started in this release. In the dashboard, Evaluations is removed from the navigation, and the page shows only a notice that evaluations are not part of this release. There is no way to list or read evaluations with an API key during the beta. The scorers below are implemented and describe what an evaluation measures.
Risicare includes 13 built-in scorers for evaluating LLM outputs across RAG, safety, agent behavior, and general quality.
Overview
As designed, an evaluation is requested through the API or the dashboard's Evaluations page, naming the scorers to run against a set of traces.
Scorer Categories
RAG Scorers
Evaluate retrieval-augmented generation quality:
| Scorer | Class | Description |
|---|---|---|
faithfulness | FaithfulnessScorer | Does the response stay faithful to retrieved context? |
answer_relevancy | AnswerRelevancyScorer | Is the response relevant to the query? |
context_precision | ContextPrecisionScorer | How precise is the retrieved context? |
context_recall | ContextRecallScorer | Does the context contain all needed information? |
hallucination | HallucinationScorer | Does the response contain hallucinated information? |
Safety Scorers
Detect harmful or inappropriate content:
| Scorer | Class | Description |
|---|---|---|
toxicity | ToxicityScorer | Offensive or harmful language |
bias | BiasScorer | Unfair or prejudiced content |
pii_leakage | PIILeakageScorer | Personal identifiable information leakage |
Agent Scorers
Evaluate agent behavior:
| Scorer | Class | Description |
|---|---|---|
tool_correctness | ToolCorrectnessScorer | Did the agent select appropriate tools? |
task_completion | TaskCompletionScorer | Did the agent complete the task? |
goal_accuracy | GoalAccuracyScorer | How accurately did the agent achieve the goal? |
General Scorers
General quality metrics:
| Scorer | Class | Description |
|---|---|---|
g_eval | GEvalScorer | General evaluation (coherence, structure, quality) |
factuality | FactualityScorer | Is the response factually correct? |
ScorerInput Fields
All scorers accept a ScorerInput dataclass. Each scorer uses a subset of these fields based on its required_fields.
| Field | Type | Description |
|---|---|---|
trace_id | str | Unique identifier for the trace being evaluated (required) |
span_id | str | None | Optional span identifier within the trace |
question | str | None | The user's question/query (RAG scorers) |
answer | str | None | The AI's response/answer to evaluate (RAG scorers) |
contexts | list[str] | List of context passages retrieved for RAG |
ground_truth | str | None | The expected/correct answer for comparison |
expected_tools | list[str] | List of tool names the agent should have used |
used_tools | list[str] | List of tool names the agent actually used |
tool_calls | list[dict] | Detailed tool call information with parameters |
task_description | str | None | Description of the task assigned to the agent |
goal | str | None | The goal the agent was trying to achieve |
output_text | str | None | Generic output text to evaluate |
input_text | str | None | Generic input text for context |
custom_criteria | str | None | User-defined evaluation criteria (G-Eval) |
evaluation_steps | list[str] | Steps to follow during evaluation (G-Eval) |
metadata | dict | Additional key-value metadata |
Required Fields by Scorer
| Scorer | Required Fields |
|---|---|
faithfulness | answer, contexts |
answer_relevancy | question, answer |
context_precision | question, contexts |
context_recall | contexts, ground_truth |
hallucination | answer, contexts |
toxicity | output_text |
bias | output_text |
pii_leakage | output_text |
tool_correctness | (none required, uses expected_tools and used_tools) |
task_completion | task_description, output_text |
goal_accuracy | goal, output_text |
g_eval | output_text |
factuality | output_text |
Requesting an Evaluation
API
Not available with an API key during the beta. As designed, an evaluation request names the evaluation, its type, the traces to evaluate and the scorers to run (criteria).
Dashboard
While evaluations are held for the beta, the dashboard's Evaluations page shows only a notice that evaluations are not part of this release, and it is removed from the navigation.
Evaluation Results
Results include:
{
"trace_id": "abc123",
"evaluations": [
{
"scorer": "faithfulness",
"score": 0.92,
"passed": true,
"reasoning": "Response accurately reflects the retrieved context..."
},
{
"scorer": "toxicity",
"score": 0.01,
"passed": true,
"reasoning": "No toxic content detected."
}
]
}Running Evaluations
init() does not take an evaluations argument — evaluations are not configured in the SDK. As designed, they are requested through the API (held for the beta, see above), naming the scorers to run in the criteria field. See the Scorers reference for the scorer list.