xberg

v2026.09.24

Extract text, tables, metadata, and images from 107 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg. Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output format), batch processing, error handling, and plugins.

GitHub
安装命令
npx skhub add xberg-io/xberg
Markdown
SKILL.md
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:a350cdaa08f8e82becd82ed51d10fc81771a4cf74c1b619e65b424101f0a612c Source-Hash: blake3:b756700854c32ac97ee5711509611d6660d4108a4d1308336f1b0a88cde1ad39 Schema-Version: v1 -->

Xberg Document Extraction

Xberg is a document intelligence library with a Rust core and bindings for Python, TypeScript/Node.js, Ruby, PHP, Go, Java, C#, Elixir, WebAssembly, Dart, Kotlin Android, Swift, Zig, and C. It extracts text, tables, metadata, and images from 107 formats across 140 unique file extensions and accepts 53 compatibility MIME aliases, including PDF, Office documents, images, HTML, email, archives, and academic formats.

Use this skill when writing code that:

  • Extracts text or metadata from documents
  • Performs OCR on scanned documents or images
  • Batch-processes multiple files
  • Configures extraction options (output format, chunking, OCR, language detection)
  • Implements custom plugins (post-processors, validators, OCR backends)

If the xberg MCP server is registered in this session, prefer its tools over shelling out to the CLI — they expose the same extraction surface with structured arguments and results.

Installation

Python

pip install xberg

Node.js

npm install @xberg-io/xberg

Rust

cargo add xberg
# Cargo.toml
[dependencies]
xberg = { version = "1.1.0", features = ["full"] }
tokio = { version = "1", features = ["full"] }
# feature flags: pdf, ocr, chunking, embeddings, language-detection, keywords, api, mcp
#                (or "formats" / "full" aggregates); tokio-runtime is on by default

CLI

brew install xberg-io/tap/xberg
# or run without a persistent install (the CLI proxy package self-installs the binary):
npx @xberg-io/xberg-cli --help
uvx --from xberg-cli xberg --help
# or download a prebuilt binary from the latest GitHub release:
#   https://github.com/xberg-io/xberg/releases/latest
# or build from source:
cargo install xberg-cli

Quick Start

The library entry points are extract(input, config) and extract_batch(inputs, config). Both return an ExtractionResult envelope — the extracted document(s) live in result.results, and per-document data (content, tables, metadata, …) is on each result.results[i]. Python and Node are async-only.

Python

import asyncio
from xberg import ExtractInput, extract, ExtractionConfig

async def main() -> None:
    result = await extract(ExtractInput(uri="document.pdf"), ExtractionConfig())
    doc = result.results[0]
    print(doc.content)    # extracted text
    print(doc.metadata)   # document metadata
    print(doc.tables)     # extracted tables

asyncio.run(main())

Node.js

import { extract } from "@xberg-io/xberg";

const output = await extract({ kind: "uri", uri: "document.pdf" });
const doc = output.results[0];
console.log(doc.content);
console.log(doc.metadata);
console.log(doc.tables);

Rust

use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result<()> {
    let output = extract(ExtractInput::from_uri("document.pdf"), &ExtractionConfig::default()).await?;
    println!("{}", output.results[0].content);
    Ok(())
}

CLI

xberg extract document.pdf
xberg extract document.pdf --format json
xberg extract document.pdf --content-format markdown

Configuration

All languages use the same configuration structure with language-appropriate naming conventions.

Python (snake_case)

from xberg import (
    ExtractInput, extract,
    ExtractionConfig, OcrConfig, TesseractConfig, PdfConfig, ChunkingConfig, OutputFormat,
)

config = ExtractionConfig(
    ocr=OcrConfig(
        backend="tesseract",
        language=["eng"],
        tesseract_config=TesseractConfig(psm=6, enable_table_detection=True),
    ),
    pdf_options=PdfConfig(passwords=["secret123"]),
    chunking=ChunkingConfig(max_characters=1000, overlap=200),
    output_format=OutputFormat("markdown"),
)

result = await extract(ExtractInput(uri="document.pdf"), config)

Node.js (camelCase)

import { extract, type ExtractionConfig } from "@xberg-io/xberg";

const config: ExtractionConfig = {
  ocr: { backend: "tesseract", language: ["eng"] },
  pdfOptions: { passwords: ["secret123"] },
  chunking: { maxCharacters: 1000, overlap: 200 },
  outputFormat: "markdown",
};

const output = await extract({ kind: "uri", uri: "document.pdf" }, config);

Rust (snake_case)

use xberg::{extract, ExtractInput, ExtractionConfig, OcrConfig, ChunkingConfig, OutputFormat};

let config = ExtractionConfig {
    ocr: Some(OcrConfig {
        backend: "tesseract".into(),
        language: vec!["eng".to_string()],
        ..Default::default()
    }),
    chunking: Some(ChunkingConfig {
        max_characters: 1000,
        overlap: 200,
        ..Default::default()
    }),
    output_format: OutputFormat::Markdown,
    ..Default::default()
};

let output = extract(ExtractInput::from_uri("document.pdf"), &config).await?;

Config File (TOML)

output_format = "markdown"

[ocr]
backend = "tesseract"
language = "eng"

[chunking]
max_characters = 1000
overlap = 200

[pdf_options]
passwords = ["secret123"]
# CLI: auto-discovers xberg.toml in current/parent directories
xberg extract doc.pdf
# or explicit:
xberg extract doc.pdf --config xberg.toml
xberg extract doc.pdf --config-json '{"ocr":{"backend":"tesseract","language":"deu"}}'

Batch Processing

extract_batch takes a list of ExtractInputs and returns one envelope whose results array holds a document per input (in input order); per-input failures are reported in result.errors.

Python

from xberg import ExtractInput, extract_batch, ExtractionConfig

inputs = [
    ExtractInput(uri="doc1.pdf"),
    ExtractInput(uri="doc2.docx"),
    ExtractInput(uri="doc3.xlsx"),
]
output = await extract_batch(inputs, ExtractionConfig())

for doc in output.results:
    print(f"{len(doc.content)} chars extracted")

Node.js

import { extractBatch } from "@xberg-io/xberg";

const output = await extractBatch([
  { kind: "uri", uri: "doc1.pdf" },
  { kind: "uri", uri: "doc2.docx" },
]);
for (const doc of output.results) {
  console.log(`${doc.content.length} chars`);
}

Rust

use xberg::{extract_batch, ExtractInput, ExtractionConfig};

let config = ExtractionConfig::default();
let inputs = vec![ExtractInput::from_uri("doc1.pdf"), ExtractInput::from_uri("doc2.docx")];
let output = extract_batch(inputs, &config).await?;

CLI

xberg batch *.pdf --format json
xberg batch docs/*.docx --content-format markdown

OCR

OCR runs automatically for images and scanned PDFs. Tesseract is the default backend (native binding, no external install required).

Backends

Select with OcrConfig.backend:

  • tesseract (default): built-in native binding. All Tesseract languages supported.
  • paddleocr ("paddleocr" / "paddle-ocr"): ONNX-based PaddleOCR.
  • vlm: Vision-Language-Model OCR (configure via OcrConfig.vlm_config).

Custom backends can be registered in Python/Node via register_ocr_backend (see Advanced Features).

Language Codes

config = ExtractionConfig(ocr=OcrConfig(language=["eng"]))          # English
config = ExtractionConfig(ocr=OcrConfig(language=["eng", "deu"]))   # Multiple
# The single-string shorthand ("eng+deu") is only accepted in config files / --config-json,
# not in the OcrConfig constructor (Python takes a list, Node takes an array).

Force OCR

config = ExtractionConfig(force_ocr=True)  # OCR even if text is extractable

Result Envelope and Document Fields

extract / extract_batch return an ExtractionResult envelope: results (list of documents), errors (per-input failures), and summary (counts). Per-document fields live on each document in results — bind doc = result.results[0] (Python/Node) or &output.results[0] (Rust) first.

FieldPython (doc.)Node.js (doc.)Rust (document.)Description
Text contentcontentcontentcontentExtracted text (str/String)
MIME typemime_typemimeTypemime_typeInput document MIME type
MetadatametadatametadatametadataDocument metadata (flat mapping)
TablestablestablestablesExtracted tables with cells + markdown
Languagesdetected_languagesdetectedLanguagesdetected_languagesDetected languages (if enabled)
ChunkschunkschunkschunksText chunks (if chunking enabled)
ImagesimagesimagesimagesExtracted images (if enabled)
ElementselementselementselementsSemantic elements (if element_based format)
PagespagespagespagesPer-page content (if page extraction enabled)
Keywordsextracted_keywordsextractedKeywordsextracted_keywordsExtracted keywords (if enabled)

Error Handling

Python

extract / extract_batch raise a plain RuntimeError on failure — the typed XbergError subclasses are not raised by these entry points, so catch RuntimeError. Per-input failures during extract_batch are reported non-fatally in result.errors.

from xberg import ExtractInput, extract, ExtractionConfig

try:
    result = await extract(ExtractInput(uri="file.pdf"), ExtractionConfig())
    for err in result.errors:
        print(f"Per-input error: {err}")
except RuntimeError as e:
    print(f"Extraction failed: {e}")

Node.js

The Node binding throws plain Error objects (it does not export typed error subclasses). Catch with instanceof Error, and inspect output.errors for non-fatal per-input failures.

import { extract } from "@xberg-io/xberg";

try {
  const output = await extract({ kind: "uri", uri: "file.pdf" });
  if (output.errors.length > 0) {
    console.error("Per-input errors:", output.errors);
  }
} catch (e) {
  if (e instanceof Error) {
    console.error(`Extraction failed: ${e.message}`);
  }
}

Rust

use xberg::{extract, ExtractInput, ExtractionConfig, XbergError};

let config = ExtractionConfig::default();
match extract(ExtractInput::from_uri("file.pdf"), &config).await {
    Ok(output) => println!("{}", output.results[0].content),
    Err(XbergError::Parsing { message, .. }) => eprintln!("Parse error: {message}"),
    Err(XbergError::Ocr { message, .. }) => eprintln!("OCR error: {message}"),
    Err(XbergError::UnsupportedFormat(mime)) => eprintln!("Unsupported: {mime}"),
    Err(e) => eprintln!("Error: {e}"),
}

Common Pitfalls

  1. Result is an envelope: extract / extract_batch return ExtractionResult with results, errors, and summary. Per-document fields (content, tables, chunks, …) are on result.results[i], NOT on the top-level return.
  2. Async-only: Python and Node have no sync variants — always await extract(...). Rust extract is async; use #[tokio::main] or an async context.
  3. Build the input: pass an ExtractInput, not a bare path. Use ExtractInput(uri=...) / ExtractInput::from_uri(...) (Python/Rust) or { kind: "uri", uri: "..." } (Node); for bytes use kind="bytes" with bytes/mime_type.
  4. Python ChunkingConfig fields: construct with max_characters and overlap (defaults 1000 / 200); these are also the readable attributes. When passing config as a dict/JSON, the max_chars / max_overlap aliases are also accepted. Node uses maxCharacters / overlap; Rust struct fields are max_characters / overlap.
  5. Python errors: extract / extract_batch raise a plain RuntimeError on failure, not typed XbergError subclasses — catch RuntimeError. Node throws plain Error (no typed error subclasses).
  6. Rust extract signature: extract(input, &config) — the config is a reference. Use &ExtractionConfig::default() for defaults.
  7. CLI --format vs --content-format: --format controls CLI output (text, json, or toon). --content-format controls content rendering (plain, markdown, djot, html, json, or doctags).
  8. Config file field names: Use snake_case in TOML/YAML/JSON config files — [chunking] fields are max_characters and overlap; other fields use names like output_format, pdf_options.

Supported Formats (Summary)

CategoryExtensions
PDF.pdf
Word.docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6, .hwp, .hwpx
Spreadsheets.xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers
Presentations.pptx, .pptm, .ppt, .pps, .ppsx, .potx, .potm, .pot, .odp, .key
eBooks.epub, .fb2
Images.png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif, .jp2, .jpg2, .j2c, .j2k, .jpc, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm, .heic, .heics, .heif, .heifs, .hif, .avif, .avcs, .svg
Markup.html, .htm, .xhtml, .xht, .xml, .kml
Data.json, .geojson, .jsonl, .ndjson, .yaml, .yml, .toml, .csv, .tsv, .dbf, .sqlite, .sqlite3, .db, .gpkg, .gpkx
Text.txt, .adoc, .asciidoc, .vtt, .md, .markdown, .commonmark, .qmd, .rmd, .mdx, .djot, .dj, .doctags, .rst, .org, .rtf
Email.eml, .msg, .pst
Archives.zip, .tar, .tgz, .gz, .7z
Audio/Video.mp3, .mpga, .m4a, .wav, .webm, .mp4, .mpg4, .mp4v, .m4v, .mpeg, .mpg, .mpe, .m1v, .m2v
Academic.bib, .ris, .nbib, .enw, .tex, .latex, .typ, .typst, .jats, .nxml, .ipynb, .docbook, .dbk, .docbook4, .docbook5, .opml

CSL JSON is supported through an explicit MIME type but does not have a registered file extension.

See references/supported-formats.md for the complete format reference with MIME types.

Additional Resources

Detailed reference files for specific topics:

Related skills

Task-focused sibling skills go deeper than this overview:

  • extracting-with-ocr — OCR backends, language packs, force-OCR, tuning.
  • extracting-tables — layout-aware table detection and table models.
  • chunking — chunk size/overlap, markdown/yaml/semantic chunkers, the chunk command.
  • extracting-keywords — YAKE/RAKE keywords, language detection, the embed command.
  • batch-extraction — the batch command, --file-configs, parallelism, error recovery.
  • picking-a-format — choosing --format / --content-format per consumer.

Full documentation: https://docs.xberg.io GitHub: https://github.com/xberg-io/xberg

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

plugin/skills/xberg

默认分支

main

最新提交

4f7e59c

Tree SHA

31b4bc5