giskard-rag

v2026.09.24

Giskard RAGET (RAG Evaluation Toolkit): automatic testset generation (simple / complex / distracting / conversational), component-level scoring (retriever / generator / rewriter), hallucination and bias tests, CI integration. Compared to RAGAS and DeepEval. USE WHEN: user mentions "Giskard", "RAGET", "Giskard RAG toolkit", "automatic testset generation", "component-level RAG scoring", "hallucination test Giskard" DO NOT USE FOR: general RAGAS usage - use `rag-evaluation`; Stanford ARES - use `ares-framework`; CI/CD wiring - use `continuous-evaluation`

GitHub
Install command
npx skhub add claude-dev-suite/giskard-rag
Markdown
SKILL.md

Giskard RAGET

What Giskard RAGET Brings

RAGET (part of the Giskard testing library) generates a diverse test set from your own documents and evaluates the RAG pipeline component by component — retriever, generator, query rewriter, and knowledge base coverage — instead of only scoring end-to-end answers.

Unique strengths:

  • Test diversity across 5 question types (simple, complex, distracting, conversational, situational).
  • Per-component diagnosis: "Is the retriever the problem or the generator?"
  • Bias and hallucination scanners inherited from Giskard core.
  • HTML report output suitable for review meetings.

Install

pip install "giskard[llm]"

Question Types

TypeWhat it stresses
simpleBaseline factual retrieval
complexMulti-hop reasoning across chunks
distractingRelevant + irrelevant-but-similar chunks
situationalUser persona / context in the prompt
conversationalMulti-turn dialogue history
doubleTwo distinct questions in one utterance

Balanced distribution catches different failure modes.

Testset Generation

from giskard.rag import KnowledgeBase, generate_testset
import pandas as pd

kb_df = pd.DataFrame({
    "text": [open(p).read() for p in doc_paths],
    "source": doc_paths,
})
kb = KnowledgeBase(kb_df)

testset = generate_testset(
    kb,
    num_questions=120,
    agent_description="Support bot for a SaaS analytics product. Users ask about"
                      " dashboards, SQL exports, billing, and API keys.",
    language="en",
)
testset.save("ragtest.jsonl")

Review the generated testset — synthetic data is never perfect. Filter out questions the KB clearly cannot answer.

Scoring a Pipeline

from giskard.rag import evaluate, QATestset

testset = QATestset.load("ragtest.jsonl")

def answer_fn(question: str, history=None) -> str:
    """Adapter to your RAG pipeline."""
    return my_rag_chain.invoke({"question": question, "history": history or []})

report = evaluate(
    answer_fn,
    testset=testset,
    knowledge_base=kb,
)

report.save("ragtest_report/")
report.to_pandas().head()

The report includes:

  • Correctness by question type and topic.
  • Component scores: retriever, generator, router, rewriter, knowledge base.
  • Failure list with reasons (context missing vs wrong generation vs both).

Component-Level Diagnostics

RAGET separates failures into buckets so you know where to invest:

BucketMeaningFix
RetrieverRelevant chunk not in top-KRe-chunk, reranker, hybrid search
GeneratorCorrect context, wrong answerPrompt edit, stronger model, few-shot
RewriterQuery reformulation lost meaningDisable or tune rewriter
RouterWrong index/tool selectedUpdate routing rules or retrain classifier
Knowledge BaseAnswer not in KB at allAdd docs; do NOT tune RAG

RAGAS aggregates everything; RAGET tells you which module to touch.

Custom LLM for Generation and Judging

from giskard.llm.client.openai import OpenAIClient
import giskard

giskard.llm.set_llm_api("openai")
giskard.llm.set_default_client(OpenAIClient(model="gpt-4o-mini"))

# Or Anthropic via LangChain wrapper
from giskard.llm.client.litellm import LiteLLMClient
giskard.llm.set_default_client(
    LiteLLMClient(model="anthropic/claude-sonnet-4-5-20250929"))

Hallucination and Bias Scan

import giskard

scan_results = giskard.scan(
    model=my_giskard_model,
    dataset=my_giskard_dataset,
    only=["hallucination", "stereotype"],
)
scan_results.to_html("scan_report.html")

Catches:

  • Ungrounded claims (hallucination).
  • Demographic stereotypes and sycophancy.
  • Prompt injection susceptibility.
  • Sensitivity to wording changes.

Comparison with RAGAS and DeepEval

AspectRAGASDeepEvalGiskard RAGET
Testset generationBasic, distribution-tunableVia Synthesizer moduleStrong, 6 question types with personas
Component-level scoringNoPartialYes
Hallucination scannerFaithfulness metricHallucination metricBuilt-in + prompt-injection scan
HTML reportNoNoYes
CI integrationLangSmith / customPytest-nativePytest via giskard.scan.run
Learning curveLowLow-mediumMedium
Best forBaseline metricsCI gatesDiagnosing where the RAG fails

Not mutually exclusive — teams often use RAGAS nightly for metrics and Giskard weekly for diagnostics.

Pytest CI Integration

# tests/test_rag_quality.py
import pytest
from giskard.rag import QATestset, evaluate

testset = QATestset.load("ragtest.jsonl")

@pytest.fixture(scope="module")
def report():
    return evaluate(my_answer_fn, testset=testset, knowledge_base=kb)

def test_overall_correctness(report):
    assert report.correctness_by_question_type()["simple"] >= 0.85

def test_retriever_recall(report):
    assert report.component_scores()["Retriever"] >= 0.80

def test_no_hallucination_regressions(report):
    failures = report.failures()
    hallucinations = [f for f in failures if f.reason == "hallucination"]
    assert len(hallucinations) <= 3

Conversational Testset Example

testset = generate_testset(
    kb,
    num_questions=60,
    question_types=["conversational", "distracting"],
    agent_description="Customer support agent with multi-turn memory.",
)

for row in testset:
    history = row.conversation_history  # list[tuple[role, content]]
    final_q = row.question
    answer = my_rag_chain.invoke({"history": history, "question": final_q})

Knowledge Base Coverage Report

from giskard.rag import KnowledgeBaseReport

kb_report = KnowledgeBaseReport(kb)
kb_report.topic_distribution()  # histogram of topics
kb_report.coverage_gaps(testset) # questions with no supporting doc

Coverage gaps indicate missing content — fix by adding docs rather than tuning RAG.

Tips That Save Hours

  • Set agent_description carefully: the generator uses it to invent realistic queries. Vague descriptions produce vague queries.
  • Seed generation for reproducibility: generate_testset(..., seed=42).
  • Regenerate the testset when the KB changes; otherwise retrieval recall drifts.
  • Start with 100-200 questions; scale to 500+ once the pipeline stabilizes.
  • Store testset in git alongside code — changes to it are reviewable.

Anti-Patterns

Anti-PatternFix
Mixing Giskard component scores across KB versionsRegenerate testset per KB version
Skipping the human review step on synthetic testsetsReview at least 30% of questions
Using default agent description "helpful assistant"Detail the domain, user persona, and scope
Running hallucination scan with a weak judge LLMUse GPT-4o or Claude Sonnet class models
Treating RAGAS faithfulness and Giskard correctness as equivalentThey measure different things; keep both
Overwriting report files without historyVersion the HTML reports for trend review

Production Checklist

  • KB dataframe built with source column for traceability
  • agent_description written specifically for the domain
  • Testset generated with balanced question-type distribution
  • At least 30% of generated questions reviewed and unusable ones filtered
  • Testset committed to git; regenerated on KB updates
  • Component-level thresholds set (retriever >= X, generator >= Y)
  • Pytest wrapper wired into CI with per-component assertions
  • Hallucination + prompt-injection scan scheduled weekly
  • HTML reports archived for trend analysis
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/rag/giskard-rag

Default branch

main

Latest commit

9496306

Tree SHA

fe4e2f1