glean-security-basics

v2026.09.24

Token security: Indexing tokens have write access -- never expose in frontend. Trigger: "glean security basics", "security-basics".

GitHub
Install command
npx skhub add jeremylongshore/glean-security-basics
Markdown
SKILL.md

Glean Security Basics

Overview

Glean indexes and searches across an enterprise's entire knowledge base — Confluence, Google Drive, Slack, GitHub, and dozens more connectors. Security concerns center on indexing token management (write-access tokens that can push content into the search index), client token scoping (user-level search permissions), and document-level access controls. A leaked indexing token allows injecting arbitrary content into enterprise search results.

API Key Management

function createGleanClient(tokenType: "indexing" | "client"): { token: string; baseUrl: string } {
  const token = tokenType === "indexing"
    ? process.env.GLEAN_INDEXING_TOKEN
    : process.env.GLEAN_CLIENT_TOKEN;
  if (!token) {
    throw new Error(`Missing GLEAN_${tokenType.toUpperCase()}_TOKEN — store in secrets manager`);
  }
  // Indexing tokens have WRITE access — never expose in frontend code
  if (tokenType === "indexing") {
    console.log("WARNING: Indexing token loaded — backend use only");
  }
  return { token, baseUrl: `https://${process.env.GLEAN_INSTANCE}.glean.com/api` };
}

Webhook Signature Verification

import crypto from "crypto";
import { Request, Response, NextFunction } from "express";

function verifyGleanWebhook(req: Request, res: Response, next: NextFunction): void {
  const signature = req.headers["x-glean-signature"] as string;
  const secret = process.env.GLEAN_WEBHOOK_SECRET!;
  const expected = crypto.createHmac("sha256", secret).update(req.body).digest("hex");
  if (!signature || !crypto.timingSafeEqual(Buffer.from(signature), Buffer.from(expected))) {
    res.status(401).send("Invalid signature");
    return;
  }
  next();
}

Input Validation

import { z } from "zod";

const IndexDocumentSchema = z.object({
  datasource: z.string().min(1).max(100),
  document_id: z.string().min(1).max(500),
  title: z.string().min(1).max(500),
  body: z.string().max(1_000_000),
  allowed_users: z.array(z.string().email()).optional(),
  allowed_groups: z.array(z.string()).optional(),
  permissions_type: z.enum(["public", "restricted", "private"]).default("restricted"),
});

function validateIndexDocument(data: unknown) {
  return IndexDocumentSchema.parse(data);
}

Data Protection

const GLEAN_SENSITIVE_FIELDS = ["indexing_token", "client_token", "document_body", "user_query", "search_results"];

function redactGleanLog(record: Record<string, unknown>): Record<string, unknown> {
  const redacted = { ...record };
  for (const field of GLEAN_SENSITIVE_FIELDS) {
    if (field in redacted) redacted[field] = "[REDACTED]";
  }
  return redacted;
}

Security Checklist

  • Indexing tokens stored server-side only, never in frontend code
  • Client tokens scoped per-user with X-Glean-Auth-Type header
  • Tokens rotated quarterly via Admin > API Tokens
  • Document permissions set via allowedUsers/allowedGroups
  • SAML SSO enforced for Glean web access
  • All API calls over HTTPS
  • Search audit logs enabled to track sensitive queries
  • Connector permissions reviewed when adding new data sources

Error Handling

VulnerabilityRiskMitigation
Leaked indexing tokenArbitrary content injected into search indexBackend-only storage + rotation
Missing document permissionsConfidential docs exposed in search resultsallowedUsers/allowedGroups on every document
Client token in frontendUser impersonation in search queriesServer-side proxy for search API
Overly broad connector scopeSensitive repos/channels indexed unintentionallyPer-connector permission review
Search queries in logsEmployee activity surveillance riskQuery redaction in logging pipeline

Prerequisites

  • A threat model identifying token custodians, untrusted inputs, source ACL authority, incident owner, and approved secret manager.
  • Separate low-privilege sandbox credentials and fictitious documents for verification; never test with a production token in a shell or CI log.
  • Rotation, revocation, and connector-disable runbooks with a named owner and tested rollback path.

Instructions

  1. Scope credentials by environment and datasource, inject them from the secret manager, and deny unknown scope or destination.
  2. Validate document schema, size, origin, and ACL before indexing; quarantine failures with opaque correlation IDs.
  3. Verify webhook authenticity before parsing, validate comparable lengths before constant-time comparison, and reject replayed or stale events.
  4. Run synthetic allow and deny probes after ACL or identity changes, logging policy revisions and aggregates rather than content.
  5. Rotate or revoke compromised credentials, disable a connector if integrity is uncertain, and preserve redacted incident evidence.

Output

Return a security receipt with environment, datasource scope, secret-reference version, validation and allow/deny outcomes, rotation/revocation state, incident correlation ID, and rollback action. Never include a token, raw webhook, query, or document body.

Examples

env=staging; source=sandbox-contracts; secret_ref=indexer-v12; signature=pass; allow_probe=pass; deny_probe=pass; rollback=connector-disabled is an auditable control result.

Resources

Next Steps

See glean-prod-checklist.

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/.curated/glean-security-basics

Default branch

main

Latest commit

e5a6c3b

Tree SHA

c2dc8e8