giskard-rag

v2026.09.24

Giskard RAGET (RAG Evaluation Toolkit): automatic testset generation (simple / complex / distracting / conversational), component-level scoring (retriever / generator / rewriter), hallucination and bias tests, CI integration. Compared to RAGAS and DeepEval. USE WHEN: user mentions "Giskard", "RAGET", "Giskard RAG toolkit", "automatic testset generation", "component-level RAG scoring", "hallucination test Giskard" DO NOT USE FOR: general RAGAS usage - use `rag-evaluation`; Stanford ARES - use `ares-framework`; CI/CD wiring - use `continuous-evaluation`

GitHub
安装命令
npx skhub add claude-dev-suite/giskard-rag
Markdown
SKILL.md

Giskard RAGET

What Giskard RAGET Brings

RAGET (part of the Giskard testing library) generates a diverse test set from your own documents and evaluates the RAG pipeline component by component — retriever, generator, query rewriter, and knowledge base coverage — instead of only scoring end-to-end answers.

Unique strengths:

  • Test diversity across 5 question types (simple, complex, distracting, conversational, situational).
  • Per-component diagnosis: "Is the retriever the problem or the generator?"
  • Bias and hallucination scanners inherited from Giskard core.
  • HTML report output suitable for review meetings.

Install

pip install "giskard[llm]"

Question Types

TypeWhat it stresses
simpleBaseline factual retrieval
complexMulti-hop reasoning across chunks
distractingRelevant + irrelevant-but-similar chunks
situationalUser persona / context in the prompt
conversationalMulti-turn dialogue history
doubleTwo distinct questions in one utterance

Balanced distribution catches different failure modes.

Testset Generation

from giskard.rag import KnowledgeBase, generate_testset
import pandas as pd

kb_df = pd.DataFrame({
    "text": [open(p).read() for p in doc_paths],
    "source": doc_paths,
})
kb = KnowledgeBase(kb_df)

testset = generate_testset(
    kb,
    num_questions=120,
    agent_description="Support bot for a SaaS analytics product. Users ask about"
                      " dashboards, SQL exports, billing, and API keys.",
    language="en",
)
testset.save("ragtest.jsonl")

Review the generated testset — synthetic data is never perfect. Filter out questions the KB clearly cannot answer.

Scoring a Pipeline

from giskard.rag import evaluate, QATestset

testset = QATestset.load("ragtest.jsonl")

def answer_fn(question: str, history=None) -> str:
    """Adapter to your RAG pipeline."""
    return my_rag_chain.invoke({"question": question, "history": history or []})

report = evaluate(
    answer_fn,
    testset=testset,
    knowledge_base=kb,
)

report.save("ragtest_report/")
report.to_pandas().head()

The report includes:

  • Correctness by question type and topic.
  • Component scores: retriever, generator, router, rewriter, knowledge base.
  • Failure list with reasons (context missing vs wrong generation vs both).

Component-Level Diagnostics

RAGET separates failures into buckets so you know where to invest:

BucketMeaningFix
RetrieverRelevant chunk not in top-KRe-chunk, reranker, hybrid search
GeneratorCorrect context, wrong answerPrompt edit, stronger model, few-shot
RewriterQuery reformulation lost meaningDisable or tune rewriter
RouterWrong index/tool selectedUpdate routing rules or retrain classifier
Knowledge BaseAnswer not in KB at allAdd docs; do NOT tune RAG

RAGAS aggregates everything; RAGET tells you which module to touch.

Custom LLM for Generation and Judging

from giskard.llm.client.openai import OpenAIClient
import giskard

giskard.llm.set_llm_api("openai")
giskard.llm.set_default_client(OpenAIClient(model="gpt-4o-mini"))

# Or Anthropic via LangChain wrapper
from giskard.llm.client.litellm import LiteLLMClient
giskard.llm.set_default_client(
    LiteLLMClient(model="anthropic/claude-sonnet-4-5-20250929"))

Hallucination and Bias Scan

import giskard

scan_results = giskard.scan(
    model=my_giskard_model,
    dataset=my_giskard_dataset,
    only=["hallucination", "stereotype"],
)
scan_results.to_html("scan_report.html")

Catches:

  • Ungrounded claims (hallucination).
  • Demographic stereotypes and sycophancy.
  • Prompt injection susceptibility.
  • Sensitivity to wording changes.

Comparison with RAGAS and DeepEval

AspectRAGASDeepEvalGiskard RAGET
Testset generationBasic, distribution-tunableVia Synthesizer moduleStrong, 6 question types with personas
Component-level scoringNoPartialYes
Hallucination scannerFaithfulness metricHallucination metricBuilt-in + prompt-injection scan
HTML reportNoNoYes
CI integrationLangSmith / customPytest-nativePytest via giskard.scan.run
Learning curveLowLow-mediumMedium
Best forBaseline metricsCI gatesDiagnosing where the RAG fails

Not mutually exclusive — teams often use RAGAS nightly for metrics and Giskard weekly for diagnostics.

Pytest CI Integration

# tests/test_rag_quality.py
import pytest
from giskard.rag import QATestset, evaluate

testset = QATestset.load("ragtest.jsonl")

@pytest.fixture(scope="module")
def report():
    return evaluate(my_answer_fn, testset=testset, knowledge_base=kb)

def test_overall_correctness(report):
    assert report.correctness_by_question_type()["simple"] >= 0.85

def test_retriever_recall(report):
    assert report.component_scores()["Retriever"] >= 0.80

def test_no_hallucination_regressions(report):
    failures = report.failures()
    hallucinations = [f for f in failures if f.reason == "hallucination"]
    assert len(hallucinations) <= 3

Conversational Testset Example

testset = generate_testset(
    kb,
    num_questions=60,
    question_types=["conversational", "distracting"],
    agent_description="Customer support agent with multi-turn memory.",
)

for row in testset:
    history = row.conversation_history  # list[tuple[role, content]]
    final_q = row.question
    answer = my_rag_chain.invoke({"history": history, "question": final_q})

Knowledge Base Coverage Report

from giskard.rag import KnowledgeBaseReport

kb_report = KnowledgeBaseReport(kb)
kb_report.topic_distribution()  # histogram of topics
kb_report.coverage_gaps(testset) # questions with no supporting doc

Coverage gaps indicate missing content — fix by adding docs rather than tuning RAG.

Tips That Save Hours

  • Set agent_description carefully: the generator uses it to invent realistic queries. Vague descriptions produce vague queries.
  • Seed generation for reproducibility: generate_testset(..., seed=42).
  • Regenerate the testset when the KB changes; otherwise retrieval recall drifts.
  • Start with 100-200 questions; scale to 500+ once the pipeline stabilizes.
  • Store testset in git alongside code — changes to it are reviewable.

Anti-Patterns

Anti-PatternFix
Mixing Giskard component scores across KB versionsRegenerate testset per KB version
Skipping the human review step on synthetic testsetsReview at least 30% of questions
Using default agent description "helpful assistant"Detail the domain, user persona, and scope
Running hallucination scan with a weak judge LLMUse GPT-4o or Claude Sonnet class models
Treating RAGAS faithfulness and Giskard correctness as equivalentThey measure different things; keep both
Overwriting report files without historyVersion the HTML reports for trend review

Production Checklist

  • KB dataframe built with source column for traceability
  • agent_description written specifically for the domain
  • Testset generated with balanced question-type distribution
  • At least 30% of generated questions reviewed and unusable ones filtered
  • Testset committed to git; regenerated on KB updates
  • Component-level thresholds set (retriever >= X, generator >= Y)
  • Pytest wrapper wired into CI with per-component assertions
  • Hallucination + prompt-injection scan scheduled weekly
  • HTML reports archived for trend analysis
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/rag/giskard-rag

默认分支

main

最新提交

9496306

Tree SHA

fe4e2f1