
Meta-Analysis & Evidence Synthesis
FreePool multiple study results into a single estimate.
Free · Opens the source repo
What Meta-Analysis & Evidence Synthesis does
The Meta-Analysis & Evidence Synthesis skill is designed for researchers and data analysts who need to combine quantitative results from multiple studies into a single pooled estimate. This skill is particularly useful when synthesizing evidence from systematic reviews, replicated experiments, or multi-cohort genome-wide association studies (GWAS). It allows users to assess the consistency of study results and quantify heterogeneity among studies, providing a comprehensive view of the data at hand.
Using this skill involves several steps, starting with the extraction of effect sizes and standard errors from the studies being analyzed. This step is critical, as it requires careful conversion of reported metrics into the format needed for pooling. The skill supports various effect size types, including odds ratios, risk ratios, hazard ratios, mean differences, and correlations. Once the data is prepared, users can choose between fixed or random effects models based on the degree of heterogeneity observed in the studies.
The skill also includes capabilities for generating a forest plot, which visually represents the effect sizes and confidence intervals for each study alongside the pooled estimate. This visual aid is essential for interpreting the results and communicating findings effectively. The Meta-Analysis & Evidence Synthesis skill is a powerful tool for anyone involved in quantitative research who needs to make sense of complex data from multiple sources, ensuring that the analysis is both rigorous and reproducible.
When to use it
Use this skill when you have results from two or more studies and need a single pooled estimate with confidence intervals.
When not to use it
Do not use this skill to find studies; it is intended for analysis after literature extraction has been performed.
What you can build with it
Pooled Analysis of Clinical Trials
A researcher combines results from several clinical trials to determine the overall effectiveness of a new medication.
Systematic Review for a Health Intervention
A team synthesizes evidence from multiple studies to evaluate the impact of a health intervention on patient outcomes.
Meta-Analysis of Genetic Studies
A geneticist pools data from multiple GWAS to identify common risk factors associated with a disease.
How to install Meta-Analysis & Evidence Synthesis
View source1. Install with the skills CLI
npx skills add mims-harvard/tooluniverse/tooluniverse-meta-analysis --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by mims-harvardMeta-Analysis & Evidence Synthesis
Pool quantitative results from multiple studies into one estimate, and judge how consistent the studies are. This is the statistical half of a systematic review (the literature-collection half is tooluniverse-literature-deep-research).
When to use this
- You have effect sizes from ≥2 studies/cohorts and want a single pooled estimate + CI.
- Synthesizing a systematic review, replicated experiments, multi-cohort GWAS, or multi-dataset associations.
- Deciding whether studies agree (low heterogeneity) or conflict (high heterogeneity).
Do NOT use it to find the studies — use tooluniverse-literature-deep-research / the literature tools for that, then bring the extracted numbers here. Before trusting any input study, consider checking it with Crossref_check_retraction.
The workflow
1. Extract (effect, SE) per study ← THE ERROR-PRONE STEP
2. Pick fixed vs random effects
3. Pool: MetaAnalysis_run
4. Read heterogeneity (I², Q, τ²)
5. Forest plot + interpret
Step 1 — Convert each study to (effect_size, se) ← do this carefully
The pooling step needs an effect size on an additive scale and its standard error. Ratio measures (OR/RR/HR) must be log-transformed first. Most reported numbers give you a CI, not an SE — derive the SE from the CI.
| What the paper reports | effect_size | se |
|---|---|---|
OR / RR / HR with 95% CI [L, U] | ln(point) | (ln(U) − ln(L)) / (2 × 1.96) |
| OR / RR / HR with a p-value (no CI) | ln(point) | ` |
GWAS / regression β with SE | β (as reported) | the reported SE |
| Two groups, means + SDs + n₁,n₂ | Hedges' g (see script) | SE of g (see script) |
Single proportion p, n | logit(p)=ln(p/(1−p)) | sqrt(1/(np) + 1/(n(1−p))) |
Pearson correlation r, n | Fisher z = atanh(r) | 1 / sqrt(n − 3) |
Critical rules
- Log-transform ratio measures. Pooling raw ORs is wrong; pool
ln(OR)and exponentiate the pooled result back. The script does this for you. - One direction. Make sure every study's effect points the same way (e.g. "exposure increases risk"); flip sign / invert the ratio for studies coded the opposite way.
- Same effect measure. Don't mix OR with HR with mean-difference in one pool.
2 × 1.96assumes a 95% CI; use2 × 1.645for 90%,2 × 2.576for 99%.
The helper script does these conversions — prefer it over hand math:
python skills/tooluniverse-meta-analysis/scripts/meta_analysis.py --input studies.csv
# studies.csv columns (use the set that matches your data):
# name, or, ci_low, ci_high (ratio + CI)
# name, beta, se (already on log/linear scale)
# name, mean1, sd1, n1, mean2, sd2, n2 (two-group means -> Hedges' g)
# name, r, n (correlation -> Fisher z)
Step 2 — Fixed vs random effects
| Use fixed-effects when | Use random-effects when |
|---|---|
| Studies estimate the same true effect (e.g. exact replications, one trial split by site) | Studies differ in population/design/dose (the usual real-world case) |
| I² is low (<25%) | I² is moderate–high, or studies are clinically heterogeneous |
When unsure, report random-effects (DerSimonian–Laird) as primary — it is the conservative default and widens the CI to reflect between-study variance.
Step 3 — Pool with MetaAnalysis_run
tu run MetaAnalysis_run '{"method":"random","studies":[
{"name":"Smith 2019","effect_size":0.41,"se":0.12},
{"name":"Lee 2021","effect_size":0.67,"se":0.18},
{"name":"Garcia 2023","effect_size":0.33,"se":0.10}]}'
Returns pooled_effect, pooled_se, pooled_ci_lower/upper, pooled_z, pooled_p_value, a heterogeneity block (Q, Q_df, Q_p_value, I_squared, tau_squared), and per_study weights + CIs.
Scale foot-gun — read this.
MetaAnalysis_runpools whatever scale you hand it and has no idea your inputs were ratios. For an OR/RR/HR you MUST pass the log-transformedeffect_size+sefrom Step 1 (e.g.ln(1.42)=0.351, not1.42) — feeding raw ratios silently produces a wrong pooled value with no error. And the values it returns — including its proseinterpretationstring — are on that same log scale. So: ignore the tool'sinterpretationfield for ratios, andexp()thepooled_effectand CI bounds back to the OR/RR/HR scale yourself before reporting. The helper script avoids all of this — it takes raw ORs, tracks the scale, and prints results already back-transformed.
Step 4 — Interpret heterogeneity (decides the story)
| I² | Heterogeneity | What it means |
|---|---|---|
| 0–25% | Low | Studies largely agree; fixed-effects is defensible |
| 25–50% | Moderate | Prefer random-effects; note the variability |
| 50–75% | Substantial | Random-effects; investigate sources (subgroup / meta-regression) |
| >75% | Considerable | Pooling may be inappropriate — explain why studies differ instead |
Q_p_value < 0.10→ statistically significant heterogeneity (Q is low-powered, so 0.10 not 0.05).tau_squaredis the between-study variance on the effect scale;> 0is what random-effects adds over fixed.- Borderline I² (≈25–50%) with a non-significant Q (
Q_p_value ≫ 0.10), especially with few studies: fixed and random-effects converge — report random-effects as primary and note that fixed-effects agrees. Don't agonize over the model choice when both give essentially the same pooled estimate.
Step 5 — Forest plot + report
The script prints a text forest plot (per-study effect, CI, weight%, and the pooled diamond). Report, in order:
- Pooled estimate + 95% CI + p (on the interpretable scale — exponentiate ratios back).
- Number of studies and total N.
- Heterogeneity: I² + Q p-value + the model you chose and why.
- Direction/consistency: do all studies point the same way?
Example: "Across 3 cohorts (N=4,210), the pooled OR was 1.51 (95% CI 1.33–1.72, p=3.4×10⁻⁶), random-effects. Heterogeneity was substantial (I²=55%, Q p=0.11), so the random-effects model is reported; all three studies showed the same direction of effect."
Honest limitations
- Garbage in, garbage out. Meta-analysis cannot fix biased primary studies; check input study quality (and retraction status via
Crossref_check_retraction) first. - Publication bias. A pooled estimate from only published studies is likely inflated. With ≥10 studies, inspect a funnel plot / Egger's test (the script notes this); with <10, state that small-study bias cannot be assessed.
- Ecological / aggregation issues. Pooling study-level summaries is not the same as pooling individual patient data.
- Don't over-pool. With I²>75% and clinically different studies, a single number can mislead — describe the variation instead.
Related skills
tooluniverse-literature-deep-research— find and grade the studies to feed in.tooluniverse-statistical-modeling— single-study regression, Cox, ORs (seereferences/cox_regression.mdfor HR extraction).tooluniverse-gwas-study-explorer/tooluniverse-gwas-finemapping— GWAS-specific multi-cohort analysis.
Frequently asked questions about Meta-Analysis & Evidence Synthesis
Similar skills
Scientific Problem Selection
Streamline your research problem selection process.
Nextflow Development
Run nf-core bioinformatics pipelines with ease.
Nature Reviewer Assessment
Simulate peer review for scientific manuscripts.
Research Writing Pipeline
Streamline your scientific writing with structured proposal-first methodologies.
Nature Literature Downloader
Efficiently download academic literature from various sources.
Auto Research
Streamline your NeMo-RL experiments with automated workflows.
