New to Claude Skills? Learn how to install them →

The Best Research Skills for AI Agents

A curated selection of the strongest research skills for AI agents, evidence mapping, literature review, trend scanning and document verification, with honest caveats.

July 31, 2026
Get Claude Skills
9 min read

What makes a good research skill

Research is a category where the honest answer to "is this any good" depends entirely on whether you can check its work. A skill that hands you a confident paragraph of synthesis is less useful than one that hands you the same synthesis with a source trail attached, because the value of automated research isn't replacing your judgment, it's giving you a faster starting point that you can still verify. The strongest skills in this list are explicit about that: they produce evidence maps, screenshot-backed claims, or structured citation reports, not just prose.

The second thing worth checking is scope. "Research" covers wildly different tasks, synthesising academic literature, scanning what people said on social media last month, verifying a manuscript's references, deciding what to work on next, and a skill built for one of those is usually a poor fit for another, even though the word "research" appears in both descriptions. Match the skill to the actual shape of your question before judging it.

This is a curated selection based on quality, popularity and maintenance activity across each skill's source repository, not a benchmarked or independently tested ranking. Every fact below (star count, licence, requirements) comes from each skill's own listing page.

The best research skills for AI agents

Ordered from most broadly applicable to most specialised, general-purpose research and evidence tools first, academic-manuscript-specific tools last.

BMad Deep Recon

A flexible research tool with three modes: drafting a deep-research prompt for you to run in an external tool like ChatGPT, Gemini, Grok or Perplexity; processing a finished research report into a distilled, cited summary other skills can consume; or running the research itself through web fan-out. Ships type packs for market, domain, technical, competitive, user-voice and academic-literature research, plus custom types via overrides.

Who it's for: anyone whose research need doesn't fit one narrow category, its multi-mode design makes it the most generally useful entry here. Caveat: its own guidance says it's not built for casual research or quick, unstructured information gathering. It's designed around rigorous, decision-grade evidence, which is more process than a five-minute question needs. Requires Python. From bmad-code-org/bmad-method, 52,000 stars, MIT licensed.

Build Evidence Map

Builds an auditable evidence map for a contested technical decision or research synthesis, preserving supporting, contradicting, qualifying and missing evidence with exact source regions, rather than collapsing disagreement into a single summarised paragraph.

Who it's for: anyone making a consequential, contestable decision (a technical choice, a proposal review) who needs the reasoning to survive scrutiny later, not just a conclusion. Caveat: its own guidance says it's not the right tool for simple fact-checking or straightforward claims that don't need extensive evidence mapping, reach for something lighter for those. Requires Node.js and shell access. From github/awesome-copilot, 38,000 stars, MIT.

Systematic Literature Review

Runs a five-phase systematic review across multiple academic papers on a topic: searches arXiv, extracts structured metadata (research questions, methodologies, key findings, limitations) in parallel across papers, and outputs a report in APA, IEEE or BibTeX format.

Who it's for: researchers, academics and students surveying a body of literature rather than reviewing one paper. Caveat: its own guidance says it's explicitly not for single-paper tasks (the skill points to a separate academic-paper-review skill for that) or for general web research that doesn't need academic rigour. Requires Python and shell access. From bytedance/deer-flow, 80,000 stars, MIT.

Eyeball

Generates a Word document where every factual claim from a document analysis includes an inline, highlighted screenshot from the original source material, so you can visually verify each claim yourself rather than taking the summary on faith.

Who it's for: researchers and analysts who need their document analysis to be independently checkable, not just plausible-sounding. Caveat: its own guidance says it's not suited to casual reading where visual proof isn't necessary, or to documents in unsupported formats. Requires Python and shell access. From github/awesome-copilot, 38,000 stars, MIT.

Last 30 Days

Researches what people are actually saying about a topic over the past month, pulling posts and engagement from Reddit, X, YouTube, TikTok, Hacker News, Polymarket, GitHub and the wider web, with a built-in "doctor" health check that flags broken or missing sources before you rely on the output.

Who it's for: anyone who needs a current-sentiment snapshot across a broad set of platforms in one pass. Caveat: its own guidance says it's not built for historical analysis, or for sources outside the eight it covers, if your platform of interest isn't on that list, the output will have a gap it won't necessarily flag. Requires Python, Node.js and shell access. From mvanhorn/last30days-skill, 58,000 stars, MIT.

Last 30 Days Research

A narrower take on the same problem: recent community and social trend research over the past month, run through a Python-based engine and output as a structured Markdown report synthesising key themes.

Who it's for: designers, marketers and researchers who want a lighter-weight version of a trend scan than the eight-source tool above. Caveat: its own guidance says it needs Python 3.12 specifically and proper credentials for the social platforms it covers. Check that before assuming it'll run as-is. Requires Python and shell access. From nexu-io/open-design, 85,000 stars, Apache-2.0.

Account Research

Researches a specific company using Common Room data, adapting its response depth to the question (a full company overview, or a targeted answer to something narrower) and defaulting to your own segments for relevance.

Who it's for: sales and business development teams preparing for a pitch or doing account-level research before an outreach push. Caveat: its own guidance says it's not built for in-depth market research or qualitative analysis that goes beyond what Common Room's data can support. No special requirements listed. From anthropics/knowledge-work-plugins, 23,000 stars, Apache-2.0.

Scientific Problem Selection

A structured framework, grounded in Fischbach & Walsh's work on decision trees in science and engineering, for scientists choosing what to work on: pitching a new project idea, troubleshooting a stuck one, or evaluating strategic risk across options.

Who it's for: working scientists navigating a genuine research-direction decision, not a technical execution problem. Caveat: its own guidance says it's not suited to users looking for technical solutions or direct coding assistance. It's about problem selection and strategy, not execution. No special requirements listed. From anthropics/knowledge-work-plugins, 23,000 stars, Apache-2.0.

Nature Reviewer Assessment

Simulates a Nature-style or general pre-submission peer review of a scientific manuscript from the referee's perspective (not an author rebuttal) producing evidence-grounded major concerns, minor comments and blocking flags. For multiple simulated reviewers, each stays mutually blind in a separate context until all reports are frozen and compared.

Who it's for: researchers who want an honest read on how a manuscript might land with reviewers before they actually submit it. Caveat: its own guidance says it's not for drafting author rebuttals, or for informal reviews that don't need journal-level rigour. No special requirements listed. From yuan1z0825/nature-skills, 34,000 stars, Apache-2.0. The narrowest entry on this list, included to show what a genuinely specialised, well-scoped research skill looks like.

How to judge a research skill's output before you rely on it

Research skills are unusually easy to over-trust, because a confident, well-formatted summary reads as authoritative whether or not it's actually correct. A short checklist helps, and it applies whether you're using a skill for a five-minute trend check or a formal literature review:

  1. Check whether the skill shows its sources, not just its conclusions. Build Evidence Map and Eyeball are both built around this principle explicitly, evidence with exact source regions, or claims backed by inline screenshots. Treat that as the standard to hold every other research skill to, even ones that don't build it in by default.
  2. Match the skill's actual data sources to your question. Last 30 Days covers eight named platforms; if your topic lives somewhere else, the absence of a result doesn't mean there's nothing to find, it means the skill didn't look there. Read the source list before trusting a "no results" outcome.
  3. Separate synthesis from verification. A skill that produces a well-written summary has done synthesis. That's not the same as verification. It hasn't confirmed the underlying claims are true, just organised them coherently. Doublecheck-style verification is a separate step, not something synthesis automatically includes.
  4. Read the skill's source before pointing it at anything sensitive. Every skill linked in this article goes to its code on GitHub rather than a re-hosted copy. For research skills that touch live company data (Account Research, for instance) that review matters more than usual, since the data being pulled may be commercially sensitive.
  5. Re-run anything that matters before you cite it externally. Trend data and social sentiment shift fast enough that a scan from a month ago may already be stale by the time you use it in a deck or a decision. Treat "last 30 days" literally. Rerun it if the decision is more than a few weeks out from when you gathered the evidence.

Common failure modes with research skills

  • Treating a synthesis as a source. Every skill on this list that produces prose summarisation should be treated as a starting point, not a citable fact on its own. Build Evidence Map and Eyeball exist specifically because that distinction matters.
  • A skill misses a platform it doesn't cover. Last 30 Days is explicit about its eight sources; if your topic lives somewhere else (a niche forum, a private Slack) expect a real gap the skill won't necessarily surface as a gap.
  • Credential or dependency errors on script-heavy skills. Last 30 Days Research needs Python 3.12 specifically; Last 30 Days needs both Python and Node.js. Check an agent's environment against a skill's stated requirements before assuming it'll just work.
  • Using an academic-scoped skill on non-academic material. Systematic Literature Review searches arXiv specifically, pointing it at a general web research question will underperform compared to a skill actually built for that.

How to install these skills

All nine skills above install through the same command. For BMad Deep Recon:

npx skills add bmad-code-org/bmad-method/bmad-deep-recon

Add --agent claude-code or --agent codex to target a specific agent, or install manually by copying the skill folder into your agent's skills directory. Full walkthroughs for installing skills in Claude Code and installing skills in Codex CLI cover both paths, and where personal versus project-scoped install locations differ by platform:

Where agent skills live on disk across Claude Code, Codex CLI, Cursor, Antigravity and other platforms, personal versus project-scoped install paths

Check the platforms directory if you're unsure whether your agent supports SKILL.md natively or needs a compatible install path.

Combining these skills into a real workflow

A realistic research pipeline might start with BMad Deep Recon to fan out a broad question, feed a promising thread into Build Evidence Map to lay out what actually supports or contradicts a specific claim, and use Eyeball when the underlying documents need to be independently verifiable by someone else on the team. Academic work follows a different chain entirely: Systematic Literature Review to survey the field, then Nature Reviewer Assessment before submission to pressure-test the manuscript.

Because skills only load their full instructions once a request matches their description, there's little downside to installing several research skills at once, the agent only activates the one that fits each specific question. If your research process is genuinely idiosyncratic (a workflow specific to your team's decision process) writing your own skill to encode it is often more reliable than adapting one of these.

Start from the research category page for the current full listing, or browse all skills for adjacent categories.

Frequently asked questions