brightdata-reference-architecture

v2026.09.24

Create a governed Bright Data collection architecture with separate control, collection, quarantine, validation, and delivery trust zones. Use when designing or reviewing a production topology. Trigger with: "architect a Bright Data system", "draw Bright Data trust boundaries", "review collection architecture".

GitHub
安装命令
npx skhub add jeremylongshore/brightdata-reference-architecture
Markdown
SKILL.md

Bright Data Governed Collection Architecture

Overview

Design collection as a policy-governed data pipeline, not a direct application call. Separate human approval and configuration from collection workers, then quarantine and validate all results before any trusted consumer or external destination receives them.

Prerequisites

  • Approved purpose, targets, fields, products, retention, recipients, and owners
  • Current proxy, Browser API, scraper, snapshot, and delivery requirements
  • Platform identity, queue, storage, policy, and observability capabilities

Instructions

Step 1: Discover the existing planes

Read design and runtime files and Grep for Bright Data credentials, zones, API endpoints, browser sessions, queues, snapshot storage, and downstream destinations. Mark every trust transition and uncontrolled shortcut.

Step 2: Define the topology

Write or Edit an architecture with these responsibilities:

Approval + workload registry -> admission controller -> bounded collection workers
                                                   -> quarantine storage
Quarantine -> schema/policy validation -> approved internal consumer
                                    \-> approved snapshot delivery destination
Telemetry <- redacted events, error classes, usage units, and decision receipts

Keep proxy and Browser API credentials in worker-specific secret bindings and REST API keys in authorized control-plane identities.

Step 3: Apply cross-cutting controls

Specify target and destination allowlists, least privilege, environment isolation, job and byte ceilings, backpressure, idempotency, schema validation, retention, deletion, redaction, audit evidence, and independent emergency stop.

Step 4: Challenge the design

Test policy denial, revoked credentials, 429, provider 5xx, browser timeout, snapshot failure or expiry, schema drift, duplicate delivery, storage exhaustion, and downstream outage. Require an owner and recovery decision for each path.

Tool Discipline

Use Read and Grep for topology discovery and evidence gathering. Use Write and Edit for diagrams, threat models, contracts, tests, and architecture decisions. This skill does not provision Bright Data resources or initiate collection.

Output

  • Control-plane and data-plane topology with trust boundaries
  • Responsibility, secret, data, and destination matrix
  • Failure analysis, rollback path, and decision record

Examples

An admission controller accepts only signed workload manifests. Product-specific workers emit raw results to quarantine, validators release an approved schema, and delivery resolves only named destinations while redacted telemetry preserves an audit trail.

Error Handling

FailureMeaningResponse
Application code can choose arbitrary targetsAdmission boundary is missingRoute requests through the workload registry
Raw collection reaches analytics directlyValidation boundary is bypassedRequire quarantine and schema promotion
Control and collection share one broad API keyIdentity blast radius is excessiveSplit identities and scopes

Resources

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/.curated/brightdata-reference-architecture

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8