glean-security-basics

v2026.09.24

Token security: Indexing tokens have write access -- never expose in frontend. Trigger: "glean security basics", "security-basics".

GitHub
安装命令
npx skhub add jeremylongshore/glean-security-basics
Markdown
SKILL.md

Glean Security Basics

Overview

Glean indexes and searches across an enterprise's entire knowledge base — Confluence, Google Drive, Slack, GitHub, and dozens more connectors. Security concerns center on indexing token management (write-access tokens that can push content into the search index), client token scoping (user-level search permissions), and document-level access controls. A leaked indexing token allows injecting arbitrary content into enterprise search results.

API Key Management

function createGleanClient(tokenType: "indexing" | "client"): { token: string; baseUrl: string } {
  const token = tokenType === "indexing"
    ? process.env.GLEAN_INDEXING_TOKEN
    : process.env.GLEAN_CLIENT_TOKEN;
  if (!token) {
    throw new Error(`Missing GLEAN_${tokenType.toUpperCase()}_TOKEN — store in secrets manager`);
  }
  // Indexing tokens have WRITE access — never expose in frontend code
  if (tokenType === "indexing") {
    console.log("WARNING: Indexing token loaded — backend use only");
  }
  return { token, baseUrl: `https://${process.env.GLEAN_INSTANCE}.glean.com/api` };
}

Webhook Signature Verification

import crypto from "crypto";
import { Request, Response, NextFunction } from "express";

function verifyGleanWebhook(req: Request, res: Response, next: NextFunction): void {
  const signature = req.headers["x-glean-signature"] as string;
  const secret = process.env.GLEAN_WEBHOOK_SECRET!;
  const expected = crypto.createHmac("sha256", secret).update(req.body).digest("hex");
  if (!signature || !crypto.timingSafeEqual(Buffer.from(signature), Buffer.from(expected))) {
    res.status(401).send("Invalid signature");
    return;
  }
  next();
}

Input Validation

import { z } from "zod";

const IndexDocumentSchema = z.object({
  datasource: z.string().min(1).max(100),
  document_id: z.string().min(1).max(500),
  title: z.string().min(1).max(500),
  body: z.string().max(1_000_000),
  allowed_users: z.array(z.string().email()).optional(),
  allowed_groups: z.array(z.string()).optional(),
  permissions_type: z.enum(["public", "restricted", "private"]).default("restricted"),
});

function validateIndexDocument(data: unknown) {
  return IndexDocumentSchema.parse(data);
}

Data Protection

const GLEAN_SENSITIVE_FIELDS = ["indexing_token", "client_token", "document_body", "user_query", "search_results"];

function redactGleanLog(record: Record<string, unknown>): Record<string, unknown> {
  const redacted = { ...record };
  for (const field of GLEAN_SENSITIVE_FIELDS) {
    if (field in redacted) redacted[field] = "[REDACTED]";
  }
  return redacted;
}

Security Checklist

  • Indexing tokens stored server-side only, never in frontend code
  • Client tokens scoped per-user with X-Glean-Auth-Type header
  • Tokens rotated quarterly via Admin > API Tokens
  • Document permissions set via allowedUsers/allowedGroups
  • SAML SSO enforced for Glean web access
  • All API calls over HTTPS
  • Search audit logs enabled to track sensitive queries
  • Connector permissions reviewed when adding new data sources

Error Handling

VulnerabilityRiskMitigation
Leaked indexing tokenArbitrary content injected into search indexBackend-only storage + rotation
Missing document permissionsConfidential docs exposed in search resultsallowedUsers/allowedGroups on every document
Client token in frontendUser impersonation in search queriesServer-side proxy for search API
Overly broad connector scopeSensitive repos/channels indexed unintentionallyPer-connector permission review
Search queries in logsEmployee activity surveillance riskQuery redaction in logging pipeline

Prerequisites

  • A threat model identifying token custodians, untrusted inputs, source ACL authority, incident owner, and approved secret manager.
  • Separate low-privilege sandbox credentials and fictitious documents for verification; never test with a production token in a shell or CI log.
  • Rotation, revocation, and connector-disable runbooks with a named owner and tested rollback path.

Instructions

  1. Scope credentials by environment and datasource, inject them from the secret manager, and deny unknown scope or destination.
  2. Validate document schema, size, origin, and ACL before indexing; quarantine failures with opaque correlation IDs.
  3. Verify webhook authenticity before parsing, validate comparable lengths before constant-time comparison, and reject replayed or stale events.
  4. Run synthetic allow and deny probes after ACL or identity changes, logging policy revisions and aggregates rather than content.
  5. Rotate or revoke compromised credentials, disable a connector if integrity is uncertain, and preserve redacted incident evidence.

Output

Return a security receipt with environment, datasource scope, secret-reference version, validation and allow/deny outcomes, rotation/revocation state, incident correlation ID, and rollback action. Never include a token, raw webhook, query, or document body.

Examples

env=staging; source=sandbox-contracts; secret_ref=indexer-v12; signature=pass; allow_probe=pass; deny_probe=pass; rollback=connector-disabled is an auditable control result.

Resources

Next Steps

See glean-prod-checklist.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/.curated/glean-security-basics

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8