New to Claude Skills? Learn how to install them →

xberg-io on GitHub

Extracting Tables

Free

Effortlessly extract tabular data from various formats.

by xberg-io8.9k stars on xberg-io/xberg
2 views
Updated Aug 10, 2026
Get this skill

Free · Opens the source repo

What Extracting Tables does

Extracting Tables is a command-line tool designed to facilitate the extraction of structured tabular data from a variety of sources, including PDFs, spreadsheets, and images. Utilizing a layout-aware table detection model, this tool can accurately identify and reconstruct tables, making it ideal for users who need to work with financial statements, scientific data, or invoices. The tool supports multiple output formats, including Markdown and JSON, allowing for flexible integration into different workflows.

The extraction process is straightforward. Users can invoke the tool via command line, specifying the source file and the desired output format. By enabling layout-aware extraction with the --layout flag, users ensure that cell boundaries are preserved, which is crucial for maintaining the integrity of the data. The tool also allows for the selection of different table reconstruction models, enabling users to optimize for accuracy or speed based on their specific needs.

For those working with spreadsheets, Extracting Tables simplifies the process by automatically converting sheets from formats like .xlsx and .csv into structured tables without requiring layout detection. This feature makes it particularly useful for data analysts and developers who need quick access to tabular data without manual intervention. Additionally, the tool provides programmatic access via Python and Node.js, allowing developers to integrate table extraction capabilities into their applications seamlessly.

While Extracting Tables is powerful, users should be aware of its limitations. Merged cells are reconstructed as repeated values, and nested tables are flattened, which may not suit all use cases. However, the tool's capabilities in accurately detecting and reconstructing tables make it a valuable asset for anyone needing to extract structured data from unstructured formats.

When to use it

Use this tool when you need to extract tables from PDFs, images, or spreadsheets and require structured output for further processing.

When not to use it

Avoid using this skill for documents with heavily merged or nested tables, as these structures may not be preserved accurately.

What you can build with it

Extracting Financial Statements

Use this tool to extract structured financial data from PDF statements, making it easier to analyze and report.

Converting Spreadsheet Data

Quickly convert .xlsx or .csv files into structured Markdown or JSON tables for use in applications or reports.

Automating Data Ingestion

Integrate the tool into your data pipeline to automate the extraction of tabular data from various document formats.

How to install Extracting Tables

View source

1. Install with the skills CLI

npx skills add xberg-io/xberg/extracting-tables --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by xberg-io
<!-- AI-RULEZ :: GENERATED FILE — DO NOT EDIT Content-Hash: blake3:7667a52a8674605a45cc61b67e7879a0104d5e86c0d82b4bde5ced9e6e3463a8 Source-Hash: blake3:5907a9cc29a5d72bbd3eaf5b820cac5133c8724895664c64fa8eafc2227716af Schema-Version: v1 -->

Extracting tables

Use this when the user wants structured tabular data — financial statements, scientific tables, invoices, spreadsheet-style PDFs. Xberg detects tables via a layout model (RT-DETR v2) and reconstructs cell structure with a configurable table model.

Basic usage

# Markdown tables embedded in the content stream
xberg extract report.pdf --layout --content-format markdown

# Structured JSON output, tables appear under result.tables
xberg extract report.pdf --layout --format json

--layout turns on layout-aware extraction; without it, tables fall back to plain text reflow and you lose cell boundaries.

Output shapes

Two surfaces, picked via --format (CLI shape) and --content-format (content rendering):

  • Markdown tables in content--content-format markdown. Tables appear inline as | col | col | blocks. Good for LLM ingestion.
  • Structured tables array--format json. Each entry has cells[][] (rows × cols), markdown (pre-rendered), page_number, bounding_box. Use this when downstream code needs exact cell access. (bounding_box is omitted when no position data is available.)

Both are populated at once when --layout is on. The tables array is always structured; the content stream switches representation.

xberg extract financials.pdf --layout --format json \
  | jq '.result.tables[] | {page: .page_number, rows: (.cells | length)}'

Table models

--layout-table-model picks the reconstruction backend:

ModelBest forNotes
tatrdense complex tables (academic, financial)Default. Heaviest, highest accuracy.
slanet_autodispatches per-table to wired/wirelessGood when table styles are mixed.
slanet_wiredtables with visible bordersFaster than tatr.
slanet_wirelesstables without borders (whitespace-separated)For invoices, simple grids.
slanet_plushybrid wired / wirelessLighter than slanet_auto.
disabledlayout detection only, no table structureUse to skip table model cost.
xberg extract bank-statement.pdf \
  --layout --layout-table-model tatr --content-format markdown

Drop --layout-confidence when the layout model misses tables (default threshold ~0.5):

xberg extract noisy-scan.pdf --layout --layout-confidence 0.3

Spreadsheets

.xlsx, .ods, .csv, .tsv are extracted by dedicated parsers — no layout model needed. Each sheet becomes a markdown table (or structured table) automatically:

xberg extract workbook.xlsx --content-format markdown
xberg extract data.csv --format json

Pass --no-cache=true only when iterating on the same file with different configs.

Config file alternative

# `output_format` in config files equals `--content-format` on the CLI.
output_format = "markdown"

[layout]
confidence_threshold = 0.5
table_model = "tatr"

Then:

xberg extract report.pdf --format json

Programmatic access

From Python, structured tables live on the document in the result envelope (result.results[0].tables):

from xberg import ExtractInput, extract, ExtractionConfig, LayoutDetectionConfig

config = ExtractionConfig(
    layout=LayoutDetectionConfig(table_model="tatr"),
    output_format="markdown",
)
result = await extract(ExtractInput(uri="report.pdf"), config)
for table in result.results[0].tables:
    print(table.markdown)        # rendered markdown
    print(table.cells[0][0])     # cell access

Node.js mirrors this (extract, output.results[0].tables, camelCase fields). See references/python-api.md and references/nodejs-api.md in the sibling xberg skill for full type signatures.

Known limitations

  • Merged cells — reconstructed as repeated values across the spanned region; the merge is not preserved as metadata.
  • Rotated tables — enable --ocr-auto-rotate true for image-based PDFs before extraction.
  • Nested tables — flattened. Detection succeeds; structural nesting is lost.
  • Multi-page tables — each page yields a separate tables[] entry. Stitch by matching column headers if needed.
  • ONNX Runtime required — layout and table models are unavailable in WASM builds and on the Android x86_64 emulator; native targets ship full support.

Common failure modes

  • Empty tables with --layout on — confidence threshold too high or table model mismatched. Drop --layout-confidence to 0.3, try --layout-table-model tatr.
  • Markdown tables look ragged — switch --layout-table-model to slanet_wired for bordered grids or slanet_wireless for invoices.
  • Slow extractiontatr is heavy. Use slanet_auto or slanet_plus as a default; reach for tatr only when accuracy matters.

See references/cli-reference.md for the full layout flag set and references/advanced-features.md for the layout pipeline internals.

Frequently asked questions about Extracting Tables

Similar skills