glean-common-errors

v2026.09.24

Diagnose and fix common Glean API errors including indexing failures, search issues, and permission problems. Trigger: "glean error", "glean not indexing", "glean search empty", "debug glean".

GitHub
安装命令
npx skhub add jeremylongshore/glean-common-errors
Markdown
SKILL.md

Glean Common Errors

Overview

Glean provides enterprise search across connected data sources with AI-powered results. API integrations involve two distinct token types (indexing vs. client) and a custom datasource model for pushing content. Common errors stem from token type mismatches, permission misconfiguration that silently hides documents from search results, and bulk indexing failures caused by duplicate upload IDs or oversized documents. Stale results are a frequent complaint -- Glean indexes asynchronously, so newly pushed documents may take 1-5 minutes to appear in search. This reference covers authentication, indexing pipeline, and search-time issues.

Prerequisites

  • A redacted correlation ID, UTC timestamp, endpoint class, and HTTP status; omit headers, real queries, and indexed bodies.
  • A scoped read-only diagnostic credential, a named owner for any connector or ACL change, and a known-good synthetic probe.

Instructions

  1. Classify the symptom by status code and operation before changing configuration.
  2. Reproduce once with the synthetic probe and capture only status, latency band, and correlation ID.
  3. Check token scope, rate-limit budget, request shape, connector freshness, and ACL watermark in that order.
  4. Make one reversible fix at a time; access-scope changes require owner approval and allow/deny probes.
  5. Escalate a redacted diagnostic bundle if the correlation class persists after rollback.

Error Reference

CodeMessageCauseFix
401UnauthorizedInvalid or expired API tokenRegenerate at Admin > Settings > API Tokens
403Wrong token typeUsing indexing token for search APIIndexing API needs indexing token; Client API needs client token with X-Glean-Auth-Type: BEARER
400uploadId already usedDuplicate bulk upload identifierGenerate a unique UUID per upload run
400document too largeDocument body exceeds 100KB limitTruncate or split content before indexing
400invalid datasourceDatasource not registeredCreate datasource first via adddatasource endpoint
400missing required fieldDocument lacks id or titleEnsure every document has both id and title fields
403Permission deniedDocument visibility restrictedSet allowAnonymousAccess: true or add user/group to permissions
429Rate limit exceededToo many API requestsImplement exponential backoff; batch indexing calls

Error Handler

interface GleanError {
  code: number;
  message: string;
  category: "auth" | "rate_limit" | "indexing" | "permission";
}

function classifyGleanError(status: number, body: string): GleanError {
  if (status === 401) {
    return { code: 401, message: body, category: "auth" };
  }
  if (status === 429) {
    return { code: 429, message: "Rate limit exceeded", category: "rate_limit" };
  }
  if (status === 403 && body.includes("permission")) {
    return { code: 403, message: body, category: "permission" };
  }
  if (status === 400) {
    return { code: 400, message: body, category: "indexing" };
  }
  return { code: status, message: body, category: "auth" };
}

Debugging Guide

Authentication Errors

Glean uses two distinct token types. Indexing tokens authenticate bulk document uploads. Client tokens authenticate search queries and require the X-Glean-Auth-Type: BEARER header. Using the wrong token type returns 403, not 401 -- check the token type first.

Rate Limit Errors

Glean enforces per-token rate limits. Indexing operations should batch documents (up to 100 per request). Search queries are rate-limited per client token. Use Retry-After header when present and implement exponential backoff starting at 2 seconds.

Validation Errors

Bulk index uploads require a unique uploadId per run -- reusing an ID silently drops the upload. Documents must include both id and title fields. Content bodies over 100KB are rejected; truncate or split large documents. New datasources must be registered via adddatasource before any documents can be indexed against them. The datasource field in each document must exactly match the registered datasource name (case-sensitive).

Error Handling

ScenarioPatternRecovery
No search results after indexingProcessing delay (1-5 min)Wait 5 minutes, then verify with direct document lookup
Stale results returnedIndex not refreshedTrigger re-index; check datasource sync schedule
Permission mismatchUser lacks document accessAdd user/group to document permissions or enable anonymous access
Bulk upload silently droppedDuplicate uploadIdAlways generate fresh UUID per upload run
Token type confusion403 on search or indexVerify correct token type for the API being called

Quick Diagnostic

# Verify client token connectivity
curl -s -o /dev/null -w "%{http_code}" \
  -H "Authorization: Bearer $GLEAN_API_KEY" \
  -H "X-Glean-Auth-Type: BEARER" \
  https://your-domain.glean.com/api/v1/search

Output

Return error class, correlation ID, datasource, probe outcome, remediation attempted, and next owner. Never include credentials, query text, result snippets, or membership data.

Examples

status=429; source=sandbox-guides; correlation=req-opaque-17; action=backoff; retry_after=60s; synthetic_probe=recovered is sufficient for a safe rate-limit handoff.

Resources

Next Steps

See glean-debug-bundle.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/.curated/glean-common-errors

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8