New to Claude Skills? Learn how to install them →

xberg-io on GitHub

Batch Extraction

Free

Efficiently extract data from multiple files simultaneously.

by xberg-io8.9k stars on xberg-io/xberg
3 views
Updated Aug 10, 2026
Get this skill

Free · Opens the source repo

What Batch Extraction does

Batch Extraction is a powerful tool designed for developers and data analysts who need to process multiple files at once. This skill allows users to run the xberg batch command, which applies a single extraction configuration across a range of documents, enabling concurrent processing and returning results in a structured format. This means that even if some files encounter errors, the extraction process continues without interruption, making it a fault-tolerant solution for batch processing tasks.

The tool supports various file formats and can handle complex extraction scenarios. Users can specify a maximum number of concurrent extractions, which is particularly useful in environments with limited memory or when using resource-intensive operations like OCR. Additionally, Batch Extraction allows for per-file configuration overrides, enabling users to customize the extraction settings for individual files while still leveraging a shared configuration for the rest.

With options for output layout and error recovery, Batch Extraction is particularly suited for tasks that require processing large volumes of documents, such as generating reports or preparing data for analysis. The output can be tailored to different formats, including JSON and markdown, depending on the user's needs. The ability to extract images and manage output directories further enhances its utility, especially for users working with presentations or mixed content types.

Overall, Batch Extraction is an essential tool for anyone looking to streamline their document processing workflows, providing flexibility and efficiency when handling multiple files simultaneously.

When to use it

Use Batch Extraction when you need to process a large number of files quickly and efficiently, especially when they share similar extraction configurations.

When not to use it

This skill may not be suitable for single-file extractions or scenarios where detailed control over individual file processing is required without shared configurations.

What you can build with it

Processing Annual Reports

Use Batch Extraction to process a directory of annual reports in PDF format, extracting key data points for analysis.

Converting Document Formats

Batch extract content from DOCX files and convert them to markdown for easier integration into a content management system.

Image Extraction from Presentations

Extract images from a batch of PowerPoint slides and save them to a specified directory for use in other applications.

How to install Batch Extraction

View source

1. Install with the skills CLI

npx skills add xberg-io/xberg/batch-extraction --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by xberg-io
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:89aa763a66dc25e9aa2849d630b288e27b1b8e6aaebf70e4ee4b58f2e3670e73 Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->

Batch extraction

Use this when processing a directory or glob of documents in one pass. xberg batch shares one extraction config across every file, runs extractions concurrently, and returns one structured array — failures on individual files do not abort the run.

Basic usage

# Glob expands to many paths; results come back as a JSON array (default)
xberg batch *.pdf

# Mixed formats, markdown content for LLM ingestion
xberg batch docs/*.docx --content-format markdown

# Recurse with the shell, then extract
xberg batch $(find ./corpus -name '*.pdf')

batch defaults to --format json (vs --format text for single extract). Each array entry is a full extraction result, so downstream code can index by position into the input path list.

xberg batch reports/*.pdf \
  | jq '.[] | {chars: (.content | length), mime: .mime_type}'

Parallelism

--max-concurrent caps how many files extract at once (default: the CPU count, capped at 8). Lower it on memory-constrained hosts or when OCR/ML models are active, since each in-flight extraction holds its own buffers. Layout-heavy batches are further limited (1 concurrent extraction for all-PDF-layout batches, 2 for mixed layout):

# Cap at 4 concurrent extractions
xberg batch scans/*.pdf --ocr true --max-concurrent 4

--max-threads additionally caps total internal threads (Rayon, ONNX intra-op, the batch semaphore) for tightly constrained environments:

xberg batch *.pdf --max-concurrent 2 --max-threads 4

Per-file config overrides

A single shared config does not always fit. --file-configs points at a JSON file mapping each path to its own override object, merged on top of the shared config for that file only:

{
  "scan.pdf": { "force_ocr": true },
  "report.pdf": { "output_format": "markdown" },
  "data.xlsx": { "output_format": "json" }
}
xberg batch scan.pdf report.pdf data.xlsx --file-configs overrides.json

Keys are file paths (matching the paths passed on the command line); values are per-file extraction config objects in snake_case, the same shape as a config file.

Output layout

For text/toon output with image extraction, --output-dir controls where referenced image files (e.g. image_0.png) are written; the directory must already exist. JSON output embeds image bytes inline and ignores --output-dir.

mkdir -p out/images
xberg batch slides/*.pptx --extract-images true --output-dir out/images --format text

Error recovery

Batch extraction is fault-tolerant per file: one unreadable or corrupt document does not stop the rest. Inspect results for partial content and surfaced errors rather than relying on the process exit code alone. Pair with --max-concurrent to avoid exhausting memory when a few large files sit in a big batch.

Shared config

Every extract flag also applies to batch (OCR, chunking, layout, content format, etc.) and is shared across all files unless a --file-configs entry overrides it:

xberg batch invoices/*.pdf \
  --layout --layout-table-model slanet_wireless \
  --content-format markdown --max-concurrent 8

A config file works too and auto-discovers from the cwd upward:

output_format = "markdown"

[ocr]
backend = "tesseract"
language = "eng"
xberg batch corpus/*.pdf --config xberg.toml

Programmatic access

From Python, extract_batch takes a list of ExtractInputs and returns one envelope whose results array holds a document per input:

from xberg import ExtractInput, extract_batch, ExtractionConfig

config = ExtractionConfig(output_format="markdown")

inputs = [ExtractInput(uri=p) for p in ["a.pdf", "b.docx", "c.xlsx"]]
output = await extract_batch(inputs, config)

for doc in output.results:
    print(len(doc.content))

Per-input overrides go on ExtractInput.config (a FileExtractionConfig). Node.js mirrors this with extractBatch; Rust uses extract_batch(inputs, &config). See references/python-api.md, references/nodejs-api.md, and references/rust-api.md in the sibling xberg skill.

MCP

When the xberg MCP server is registered, prefer the extract_batch tool over shelling out — it takes an array of input objects and a config object and returns structured results directly.

Common pitfalls

  • Default format differsbatch defaults to --format json, extract to --format text. Set --format explicitly if a script depends on one shape.
  • --output-dir must exist — the CLI does not create it.
  • Memory blowups — large batches with OCR/layout active need a lower --max-concurrent; the default is the CPU count, capped at 8.
  • --file-configs path keys — must match the paths as passed on the command line, not absolute-resolved variants.

See references/cli-reference.md for the full batch flag set.

Frequently asked questions about Batch Extraction

Similar skills