New to Claude Skills? Learn how to install them →

agricidaniel on GitHub

Wiki Retrieve

Free

Efficiently retrieve and rank passages from your vault.

Get this skill

Free · Opens the source repo

What Wiki Retrieve does

Wiki Retrieve is a powerful tool designed for developers and researchers who need to efficiently extract and rank relevant passages from a local vault of notes. By utilizing a BM25 retrieval index, this skill allows users to perform contextual searches that can yield precise results based on the content of their vault. The process begins with the contextual-prefix.py script, which organizes text into manageable chunks while preserving important prefixes for context. This is followed by the creation of a BM25 index using bm25-index.py, enabling fast and effective retrieval of relevant information.

The retrieval process is initiated through retrieve.py, which selects candidates from the BM25 index and can optionally rerank them using a multilingual model for enhanced accuracy. This flexibility allows users to choose between a straightforward BM25 retrieval or a more nuanced semantic search, depending on their needs. The system is designed with privacy in mind, ensuring that data remains local unless explicit consent is given for external processing. This makes it suitable for sensitive information where data privacy is a concern.

Wiki Retrieve is ideal for anyone who frequently queries large sets of notes or documents and requires a reliable way to find relevant information quickly. Whether you're a developer looking to streamline your documentation process or a researcher needing to sift through extensive literature, this skill provides the tools necessary to enhance your workflow. With its focus on contextualized retrieval and privacy, it stands out as a valuable addition to any knowledge management system.

When to use it

Use this tool when you need to extract specific information from a large set of documents or notes stored locally.

When not to use it

This skill may not be suitable for real-time collaborative environments or when searching across multiple external databases.

What you can build with it

Academic Research

Use Wiki Retrieve to quickly find relevant passages from your notes when writing papers or conducting literature reviews.

Documentation Management

Developers can leverage this tool to efficiently search through project documentation and code comments, improving productivity.

Personal Knowledge Base

Enhance your personal note-taking system by retrieving specific information from a large collection of notes.

How to install Wiki Retrieve

View source

1. Install with the skills CLI

npx skills add agricidaniel/claude-obsidian/wiki-retrieve --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by agricidaniel

Retrieve relevant passages

This extension derives search data from wiki/ into .vault-meta/. It never changes canonical notes. Always pass the selected vault explicitly.

Resolve the installed product root from this skill's own location, not from the vault or current working directory:

PRODUCT_ROOT=/absolute/path/to/installed/claude-obsidian
PREFIX="$PRODUCT_ROOT/scripts/contextual-prefix.py"
BM25="$PRODUCT_ROOT/scripts/bm25-index.py"
RETRIEVE="$PRODUCT_ROOT/scripts/retrieve.py"
RERANK="$PRODUCT_ROOT/scripts/rerank.py"
test -f "$PREFIX" && test -f "$BM25" && test -f "$RETRIEVE" && test -f "$RERANK"

Pipeline

  1. contextual-prefix.py splits pages on paragraph boundaries and stores the raw chunk plus a short page-level prefix.
  2. bm25-index.py builds a local, standard-library BM25 index over the contextualized text.
  3. retrieve.py selects BM25 candidates, optionally reranks them, rejects invalid records, deduplicates by page, and returns paths and snippets.
  4. The caller reads the returned pages and performs synthesis; retrieval output is not itself evidence.

Provision locally

Preview first, then build synthetic prefixes without network egress:

python3 "$PREFIX" --vault "$VAULT" --all --no-llm --peek
python3 "$PREFIX" --vault "$VAULT" --all --no-llm
python3 "$BM25" --vault "$VAULT" build
python3 "$RETRIEVE" --vault "$VAULT" "wiki" --top 1 --no-rerank --explain

Chunk and index files are disposable runtime state. Incremental prefixing skips records whose chunk and page hashes still match. A complete scan removes surplus records for deleted pages, and the prefixer invalidates the BM25 index before changing its chunk set so a mixed stale index is not served. Prefix and BM25 build operations share the vault-wide mutation lock with every other writer; a busy vault fails closed instead of publishing a partial index.

Contextual-prefix privacy

Synthetic prefixes use only local frontmatter and page text. The Anthropic API and claude subprocess tiers can send page bodies off-machine and therefore require the user's explicit consent plus --allow-egress. Never infer consent from an API key or installed binary. Preview the scope first and state which provider will receive what data.

Remote Ollama endpoints also require explicit approval and --allow-remote-ollama; the default reranker accepts localhost only.

Query

For a strictly read-only lookup, use the prebuilt BM25 index:

python3 "$RETRIEVE" --vault "$VAULT" "$QUERY" --top 5 --no-rerank --explain

For an explicitly requested rerank, omit --no-rerank. The default is Ollama's multilingual nomic-embed-text-v2-moe model (approximately 958 MB); the product never pulls it automatically. To use an already-installed, smaller, English-oriented v1.5 model, pass --model nomic-embed-text explicitly. Nomic models use search_query: for the query and search_document: for candidate text. Nomic v2 has a 512-token input context and Ollama truncates longer embedding inputs by default; BM25 still scores the complete chunk. Embeddings are cached by exact model, input scheme, and hash of the exact prefixed input. A missing local Ollama service, missing selected model, unusable vector, or any candidate embedding failure falls back for the complete result set to the original BM25 order; it never mixes cosine and BM25 score scales.

Query input is bounded at 8,000 normalized characters and result counts must be between 1 and 1,000. Oversized queries and invalid limits fail with an actionable usage error instead of looking like an empty successful search. An untagged model request matches only the installed untagged name or its :latest alias; select any other tag explicitly.

Use direct diagnostics when needed:

python3 "$BM25" --vault "$VAULT" stats
python3 "$BM25" --vault "$VAULT" query "$QUERY" --top 10
python3 "$RERANK" --vault "$VAULT" "$QUERY" --peek
python3 "$RERANK" --vault "$VAULT" "$QUERY" --model nomic-embed-text --peek

Integrity rules

  • Accept only relative chunk and page paths whose resolved targets remain under $VAULT/.vault-meta/chunks/ and $VAULT/wiki/ respectively.
  • Reject hashless legacy chunk records and require chunk-body, page, and index hashes to match before a cached record can be built or served.
  • Reject absolute paths, symlink escapes, missing pages, mismatched chunk IDs, changed page hashes, and stale index/chunk hash pairs.
  • Rerank the full candidate set, then deduplicate by page, then apply --top.
  • An empty index is an honest no-result state. A missing or corrupt index makes retrieve.py exit 10 with a stable rebuild command; callers fall back to the standard vault query/text-search path and do not fabricate matches.
  • Do not cite benchmark percentages unless a reproducible vault-specific benchmark produced them.

Checkpoint

Observe cache readiness and privacy boundaries, think about whether lexical or semantic ranking is needed, verify returned paths and source freshness, and grow by measuring retrieval misses against a maintained local query set.

Frequently asked questions about Wiki Retrieve

Similar skills