New to Claude Skills? Learn how to install them →

trailofbits on GitHub

YARA-X Rule Authoring

Free

Create effective malware detection rules with YARA-X.

Get this skill

Free · Opens the source repo

What YARA-X Rule Authoring does

The YARA-X Rule Authoring skill provides developers and security analysts with a structured approach to writing high-quality detection rules for malware identification using YARA-X, the advanced successor to legacy YARA. This skill emphasizes the importance of crafting rules that minimize false positives while maximizing detection accuracy. By following best practices outlined in the skill, users can ensure their rules are efficient and effective, which is crucial in the fast-paced environment of threat hunting and malware analysis.

The core principles of this skill guide users through the intricacies of string selection, rule naming conventions, and performance optimization. It encourages targeting specific malware families rather than broad categories, ensuring that rules are precise and actionable. Moreover, the skill highlights the necessity of testing rules against known clean files to avoid false alarms, which can waste valuable time and resources. With YARA-X's improved performance capabilities, users can expect faster and more reliable results when deploying their detection rules.

This skill is particularly useful for those involved in malware detection and analysis, including security researchers, incident responders, and threat hunters. Whether you are writing new rules, reviewing existing ones for performance, or migrating from legacy YARA, this skill provides the necessary tools and guidance to enhance your rule authoring process. Additionally, it offers insights into analyzing specific file types, such as Chrome extensions and Android applications, using the new crx and dex modules, thus broadening the scope of malware detection capabilities.

In summary, the YARA-X Rule Authoring skill is an essential resource for anyone looking to improve their malware detection strategies through the effective use of YARA-X. By adhering to the guidelines and principles provided, users can develop robust detection rules that stand the test of time against evolving malware threats.

When to use it

Use this skill when developing new YARA-X rules for malware detection or optimizing existing ones for better performance.

When not to use it

This skill is not suitable for static analysis requiring disassembly or dynamic malware analysis; consider using dedicated tools for those purposes.

What you can build with it

Writing New Detection Rules

Utilize this skill to create new YARA-X rules tailored for specific malware threats, ensuring they are optimized for performance.

Reviewing Existing Rules

Employ the guidelines to assess and enhance the quality and efficiency of your current YARA rules.

Migrating from Legacy YARA

Follow the provided instructions to transition your existing YARA rules to the more efficient YARA-X format.

How to install YARA-X Rule Authoring

View source

1. Install with the skills CLI

npx skills add trailofbits/skills/yara-rule-authoring --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by trailofbits

YARA-X Rule Authoring

Write detection rules that catch malware without drowning in false positives.

This skill targets YARA-X, the Rust-based successor to legacy YARA — 5-10x faster regex, better errors, built-in formatter, stricter validation, new modules (crx, dex), 99% rule compatibility. It powers VirusTotal's production systems. Install with brew install yara-x or cargo install yara-x; the CLI is yr. See Migrating from Legacy YARA for existing rules.

Core Principles

  1. Strings must generate good atoms — YARA extracts 4-byte subsequences for fast matching. Strings with repeated bytes, common sequences, or under 4 bytes force slow bytecode verification on too many files.

  2. Target specific families, not categories — "Detects ransomware" catches everything and nothing. "Detects LockBit 3.0 configuration extraction routine" catches what you want.

  3. Test against goodware before deployment — A rule that fires on Windows system files is useless. Validate against VirusTotal's goodware corpus or your own clean file set.

  4. Short-circuit with cheap checks firstfilesize (instant), then magic bytes (nearly instant), then strings (cheap), then modules (expensive).

  5. Metadata is documentation — Future you (and your team) need to know what this catches, why, and where the sample came from.

When to Use

  • Writing new YARA-X rules for malware detection
  • Reviewing existing rules for quality or performance issues
  • Optimizing slow-running rulesets
  • Converting IOCs or threat intel into detection signatures
  • Debugging false positive issues
  • Preparing rules for production deployment
  • Migrating legacy YARA rules to YARA-X
  • Analyzing Chrome extensions (crx module) or Android apps (dex module)

When NOT to Use

  • Static analysis requiring disassembly → use Ghidra/IDA skills
  • Dynamic malware analysis → use sandbox analysis skills
  • Network-based detection → use Suricata/Snort skills
  • Memory forensics with Volatility → use memory forensics skills
  • Simple hash-based detection → just use hash lists

Platform Considerations

YARA works on any file type. Adapt patterns to your target:

PlatformMagic BytesBad StringsGood Strings
Windows PEuint16(0) == 0x5A4DAPI names, Windows pathsMutex names, PDB paths
macOS Mach-Ouint32(0) == 0xFEEDFACE (32-bit), 0xFEEDFACF (64-bit), uint32be(0) == 0xCAFEBABE (universal)Common Obj-C methodsKeylogger strings, persistence paths
JavaScript/Node(none needed)require, fetch, axiosObfuscator signatures, eval+decode chains
npm/pip packages(none needed)postinstall, dependenciesSuspicious package names, exfil URLs
Office docsuint32(0) == 0x04034B50VBA keywordsMacro auto-exec, encoded payloads
VS Code extensions(none needed)vscode.workspaceUncommon activationEvents, hidden file access
Chrome extensionsUse crx moduleCommon Chrome APIsPermission abuse, manifest anomalies
Android appsUse dex moduleStandard DEX structureObfuscated classes, suspicious permissions

uintNN() reads little-endian. Write the constant as the bytes reversed, or use uintNNbe() and write them in file order. A ZIP/OOXML file starts with bytes 50 4B 03 04, so it is uint32(0) == 0x04034B50uint32(0) == 0x504B0304 compiles cleanly and never matches anything. The same trap catches Mach-O universal binaries: on disk they are CA FE BA BE, so uint32(0) == 0xCAFEBABE is a dead branch; write uint32be(0) == 0xCAFEBABE or uint32(0) == 0xBEBAFECA. Verify with yr scan against one known-good sample before trusting any magic-byte check.

macOS Malware Detection

No dedicated Mach-O module exists yet — use magic bytes plus string patterns. Good indicators:

  • Keylogger artifacts: CGEventTapCreate, kCGEventKeyDown
  • SSH tunnel strings: ssh -D, tunnel, socks
  • Persistence paths: ~/Library/LaunchAgents, /Library/LaunchDaemons
  • Credential theft: security find-generic-password, keychain
// Pattern from Airbnb BinaryAlert
rule SUSP_Mac_ProtonRAT
{
    strings:
        $lib1 = "SRWebSocket" ascii          // Library indicators
        $lib2 = "SocketRocket" ascii
        $behav1 = "SSH tunnel not launched" ascii   // Behavioral indicators
        $behav2 = "Keylogger" ascii
    condition:
        (uint32(0) == 0xFEEDFACF or uint32be(0) == 0xCAFEBABE) and
        any of ($lib*) and any of ($behav*)
}

JavaScript Detection

TargetApproach
npm packagepackage.json patterns, postinstall/preinstall hooks, exfil combination: fetch + env access + credential paths
Chrome extensioncrx module
Other extensionManifest patterns, background script behaviors
Standalone JSObfuscation markers (eval+atob, fromCharCode chains), unique function/variable names, packed payloads
Minified/webpack bundleUnique strings that survive bundling (URLs, magic values); avoid function names — they get mangled

Good JS strings: Ethereum function selectors — { a9 05 9c bb } (transfer(address,uint256)), { 70 a0 82 31 } (balanceOf(address)); zero-width characters for steganography — { E2 80 8B E2 80 8C }; obfuscator signatures — _0x, var _0x; specific C2 domains and webhook URLs.

Bad JS strings: require, fetch, axios (too common); Buffer, crypto (legitimate uses everywhere); process.env alone (need specific env var names).

String Selection

Value ranking: mutex names are gold, C2 paths silver, error messages bronze. Stack strings are almost always unique. If you need more than 6 strings, you're over-fitting.

Reject a candidate string when any of these holds:

TestWhy it failsDo instead
Under 4 bytesNo atomFind a longer string
Repeated bytes (0000, 9090)Weak atomAdd surrounding context
API name (VirtualAlloc, CreateRemoteThread)Every packer and installer calls itHex pattern of the call site plus a unique marker
Appears in Windows system filesGuaranteed FPsFind something family-specific
Common path (C:\Windows\, cmd.exe)UbiquitousFind malware-specific paths
Appears in other malware familiesNot identifying this familyCombine with a family-specific marker

Everything left — unique to this family — is what the rule should rest on.

Choosing a String Type

NeedUse
Exact ASCII/Unicode text$s = "MutexName" ascii wide
Specific byte sequence$h = { 4D 5A 90 00 }
Byte sequence with variationHex wildcards: { 4D 5A ?? ?? 50 45 }
Pattern with structure (URLs, paths)Bounded regex: /https:\/\/[a-z]{5,20}\.onion/
Unknown encoding (XOR, base64)Modifier: $s = "config" xor(0x00-0xFF)

Modifier discipline: never use nocase or wide speculatively — only with confirmed evidence that case or encoding varies across samples. nocase doubles atom generation; wide doubles string matching. "If you don't have a clear reason for using those modifiers, don't do it" — Kaspersky Applied YARA.

Condition Design

Order for short-circuit: filesize <, magic bytes, strings, modules. If the condition runs past 5 lines, split into multiple rules.

all of vs any of

SituationUse
Strings are individually unique to the malwareany of them — each alone is suspicious
Strings are common but the combination is suspiciousall of them — require the full pattern
Strings have different confidence levelsGroup: all of ($core_*) and any of ($variant_*)
Seeing false positivesTighten: anyall, add more required strings

Lesson from production: rules using any of ($network_*) where the strings included fetch, axios, and http matched virtually all web applications. Switching to require a credential path AND a network call AND an exfil destination eliminated the FPs.

Grouping by Confidence

Different indicator types carry different weight — a C2 domain might be definitive while library imports need corroboration. Grouping by prefix lets you express graduated requirements:

strings:
    $a1 = "SRWebSocket" ascii            // Category A: library indicators
    $a2 = "SocketRocket" ascii
    $b1 = "SSH tunnel" ascii             // Category B: behavioral
    $b2 = "keylogger" ascii nocase
    $c1 = /https:\/\/[a-z0-9]{8,16}\.onion/   // Category C: C2

condition:
    filesize < 10MB and
    any of ($a*) and any of ($b*)        // Evidence from BOTH categories

Modules vs Byte Checks

NeedUse
imphash, rich header, authenticodePE module — too complex to replicate
Magic bytes or simple offsetsuint16/uint32 — faster, no module overhead
Section names/sizesPE module, but put the magic-byte filter FIRST
Chrome extension permissionscrx module — string parsing is fragile
LNK target pathslnk module — the format is complex

"Avoid the magic module — use explicit hex checks instead" — Neo23x0. Generalize it: if uint32() can do the job, don't load a module.

Performance

  • Regex must be anchored to a 4+ byte literal. Without one it evaluates at every file offset — catastrophic. Write /mshta\.exe http:\/\/.../, not /http:\/\/.../. If you can't anchor, use a hex pattern with wildcards.
  • Bound every regex quantifier.{0,30}, never .*. Unbounded regex is both a performance disaster and a memory explosion.
  • Bound loops with filesizefilesize < 100KB and for all i in (1..#a) : .... Unbounded #a can reach thousands in large files.
  • Prefer hex over regex where the bytes are fixed.

Before Writing: Is the Sample Packed?

SignalWhat to do
Entropy > 7.0Likely packed — find the unpacked layer first
Few or no readable stringsLikely packed — use entropy, PE structure, or packer signatures
UPX/MPRESS/custom packer detectedTarget the unpacked payload OR detect the packer itself
Readable strings availableProceed with string-based detection

Don't write rules against packed layers. The packing changes; the payload doesn't.

When Strings Fail, Pivot to Structure

If extraction returns only API names and generic paths:

Available signalUse
High entropy sectionsmath.entropy() on specific sections
Unusual import patternpe.imphash() for import-hash clustering
PE structure anomaliesSection names, sizes, characteristics
Metadata presentVersion info, timestamps, resources
Nothing uniqueThis sample may not be detectable with YARA alone

"One can try to use other file properties, such as metadata, entropy, import hashes or other data which stays constant." — Kaspersky Applied YARA Training

Debugging False Positives

  1. Which string matched?yr scan -s rule.yar false_positive.exe
  2. In a legitimate library? — add a not $fp_vendor_string exclusion
  3. A common development pattern? — replace the string with something more specific
  4. Multiple generic strings matching together? — tighten to require all, plus a unique marker
  5. Malware using a common technique? — target its specific implementation details, not the technique

When to Abandon the Approach

  • Extraction returns only API names and pathspivot to structure
  • Can't find 3 unique strings → probably packed; target the unpacked version or detect the packer
  • Rule matches goodware → 1-2 matches: investigate and tighten; 3-5: find different indicators; 6+: start over
  • Performance is terrible after optimization → architecture problem; split into focused rules or add strict pre-filters
  • The description is hard to write → the rule is too vague. If you can't explain what it catches, it catches too much

Rationalizations to Reject

When you catch yourself thinking these, stop and reconsider.

RationalizationExpert Response
"This generic string is unique enough" / "This hex pattern is unique"Unique in one sample ≠ unique across the ecosystem. Test against goodware; your intuition is wrong.
"yarGen gave me these strings"yarGen suggests, you validate. Check each one manually — expect to discard 80%.
"It works on my 10 samples"10 samples ≠ production. Use a goodware corpus.
"One rule to catch all variants"Causes FP floods. Target specific families.
"I'll make it more specific if we get FPs" / "I'll add more conditions later"Write tight rules upfront. A weak rule deployed is damage done, and FPs burn trust.
"This is just for hunting"Hunting rules become detection rules. Same quality bar.
"The API name makes it malicious"Legitimate software uses the same APIs. Need behavioral context.
"any of them is fine for these common strings"Common strings + any = FP flood. Use any of only for individually unique strings.
"This regex is specific enough"/fetch.*token/ matches all auth code. Add an exfil destination requirement.
"I'll use .* for flexibility"Unbounded regex = performance disaster plus memory explosion. Use .{0,30}.
"The JavaScript looks clean"Attackers poison legitimate code with injects. Check for eval+decode chains.
"Performance doesn't matter"One slow rule slows the entire ruleset. Optimize atoms.
"I'll use --relaxed-re-syntax everywhere"Masks real bugs. Fix the regex instead of hiding the problem.
"PEiD rules still work"Obsolete. 32-bit packers aren't relevant.

Toolkit

ToolPurpose
yr CLIyr check (validate), yr fmt (format), yr scan -s (scan, show strings), yr dump -m pe (inspect structure)
yarGenExtract candidate strings: yarGen.py -m samples/ --excludegood
FLOSSExtract obfuscated/stack strings: floss sample.exe — when yarGen comes up empty
signature-baseStudy quality examples
YARA-CIGoodware corpus testing before deployment

Master these five. Don't get distracted by tool catalogs.

Development cycle:

yr check rule.yar                                   # syntax, with precise line numbers
yr fmt -w rule.yar                                  # standardize formatting
yr dump -m pe sample.exe --output-format yaml       # inspect structure, no dummy rule needed
time yr scan -s rule.yar corpus/                    # scan with timing

Reach for yr dump when investigating which module fields are available, debugging why a module condition isn't matching, or exploring a new module (crx, lnk, dotnet) before writing against it. YARA-X error messages carry precise source locations — if yr check says line 15, the problem is on line 15.

Version-gated features: private $helper = "pattern" matches but stays out of output (v1.3.0+); // suppress: slow_pattern silences a specific warning inline (v1.4.0+); filesize < 10_000_000 numeric underscores (v1.5.0+). $_unused also suppresses unused-string warnings.

Chrome Extension Analysis (crx module)

Requires YARA-X v1.5.0+, or v1.11.0+ for permhash().

Key APIs: crx.is_crx, crx.permissions, crx.permhash() Red flags: nativeMessaging + downloads, debugger permission, content scripts on <all_urls>

import "crx"

rule SUSP_CRX_HighRiskPerms {
    condition:
        crx.is_crx and
        for any perm in crx.permissions : (perm == "debugger")
}

See crx-module.md for the full API, permission risk assessment, and example rules.

Android DEX Analysis (dex module)

Requires YARA-X v1.11.0+. Not compatible with legacy YARA's dex module — the API is completely different.

Key APIs: dex.is_dex, dex.contains_class(), dex.contains_method(), dex.contains_string() Red flags: single-letter class names (obfuscation), DexClassLoader reflection, encrypted assets

import "dex"

rule SUSP_DEX_DynamicLoading {
    condition:
        dex.is_dex and
        dex.contains_class("Ldalvik/system/DexClassLoader;")
}

See dex-module.md for the full API, obfuscation detection, and example rules.

Migrating from Legacy YARA

99% rule compatibility, but stricter validation:

yr check --relaxed-re-syntax rules/   # identify issues
# fix each one, then verify without relaxed mode:
yr check rules/
IssueLegacyYARA-X Fix
Literal { in regex/{//\{/
Invalid escapes\R silently literal\\R or R
Base64 stringsAny length3+ chars required
Negative indexing@a[-1]@a[#a - 1]
Duplicate modifiersAllowedRemove duplicates

--relaxed-re-syntax is a diagnostic, not a destination. Fix the regex.

Naming and Metadata

{CATEGORY}_{PLATFORM}_{FAMILY}_{VARIANT}_{DATE}      e.g. MAL_Win_Emotet_Loader_Jan25

Categories: MAL_ (malware), HKTL_ (hacking tool), WEBSHELL_, EXPL_, SUSP_ (suspicious), GEN_ (generic). Platforms: Win_, Lnx_, Mac_, Android_, CRX_.

Every rule needs description (starting with "Detects"), author, reference, and date:

meta:
    description = "Detects Example malware via unique mutex and C2 path"
    author = "Your Name <email@example.com>"
    reference = "https://example.com/analysis"
    date = "2025-01-29"

See style-guide.md for full conventions.

Workflow

  1. Gather samples — multiple; single-sample rules are brittle
  2. Extract candidatesyarGen -m samples/ --excludegood
  3. Validate quality — apply the string selection tests; expect to discard 80% of yarGen output
  4. Write the rule — proper metadata, cheap checks first
  5. Lint and testyr check, yr fmt, the linter script
  6. Goodware validation — VirusTotal corpus or local clean files
  7. Deploy — full metadata, then monitor for FPs

Quality signals along the way: a rule matching under 50% of known variants is too narrow; one matching goodware is too broad.

Reviewing a rule someone else wrote — run both scripts before reading the rule by eye, and quote the codes they emit:

uv run {baseDir}/scripts/yara_lint.py suspect.yar      # style, metadata, YARA-X compatibility
uv run {baseDir}/scripts/atom_analyzer.py suspect.yar  # atom quality per string

They catch the mechanical faults — short strings, FP-prone substrings, unbounded quantifiers, expensive terms ahead of cheap ones — so your attention goes to the judgement calls they cannot make: whether the strings identify this family, and whether the condition can fire on generic strings alone. Report findings by code (E002, W009) so the author can look each one up in style-guide.md.

See testing.md for the validation workflow and rule-development.md for the full step-by-step guide.

Common Mistakes

MistakeBadGood
API names as indicators"VirtualAlloc"Hex pattern of call site + unique mutex
Unbounded regex/https?:\/\/.*//https?:\/\/[a-z0-9]{8,12}\.onion/
Missing file type filterpe.imports(...) firstuint16(0) == 0x5A4D and filesize < 10MB first
Short strings"abc" (3 bytes)"abcdef" (4+ bytes)
Unescaped braces (YARA-X)/config{key}//config\{key\}/
Wrong-endian magic bytesuint32(0) == 0xCAFEBABEuint32be(0) == 0xCAFEBABE

Quality Checklist

Before deploying any rule:

  • Name follows {CATEGORY}_{PLATFORM}_{FAMILY}_{VARIANT}_{DATE}
  • Description starts with "Detects" and explains what/how
  • All required metadata present (author, reference, date)
  • Strings are unique — not API names, common paths, or format strings
  • All strings 4+ bytes with good atom potential
  • Base64 modifier only on strings with 3+ characters
  • Regex bounded, anchored to a literal, with { escaped
  • Condition starts with cheap checks (filesize, magic bytes)
  • Magic-byte constants verified against a known-good sample
  • Rule matches all target samples
  • Rule produces zero matches on the goodware corpus
  • yr check and yr fmt --check pass
  • Linter passes with no errors
  • Peer review completed

Scripts

uv run {baseDir}/scripts/yara_lint.py rule.yar      # validate style/metadata
uv run {baseDir}/scripts/atom_analyzer.py rule.yar  # check string quality

See README.md for detailed script documentation.

Further Reading

TopicDocument
Naming and metadata conventionsstyle-guide.md
Performance and atom optimizationperformance.md
String types and judgmentstrings.md
Testing and validationtesting.md
Chrome extension module (crx)crx-module.md
Android DEX module (dex)dex-module.md
Complete rule development processrule-development.md

The examples/ directory holds real, attributed rules worth reading before writing your own:

ExampleDemonstratesSource
MAL_Win_Remcos_Jan25.yarPE malware: graduated string counts, multiple rules per familyElastic Security
MAL_Mac_ProtonRAT_Jan25.yarmacOS: Mach-O magic bytes, multi-category groupingAirbnb BinaryAlert
MAL_NPM_SupplyChain_Jan25.yarnpm supply chain: real attack patterns, ERC-20 selectorsStairwell Research
SUSP_JS_Obfuscation_Jan25.yarJavaScript: obfuscator detection, density-based matchingimp0rtp3, Nils Kuhnert
SUSP_CRX_SuspiciousPermissions.yarChrome extensions: crx module, permissionsEducational

Rule repositories to learn from: Neo23x0/signature-base (17,000+ production rules), elastic/protections-artifacts (endpoint-tested), imp0rtp3/js-yara-rules (JavaScript), InQuest/awesome-yara (curated index).

Guides: YARA Style Guide and YARA Performance Guidelines (Neo23x0), YARA-X documentation.

macOS specifics: Apple's own production rules ship at /System/Library/CoreServices/XProtect.bundle/; objective-see publishes macOS malware research and samples.

Frequently asked questions about YARA-X Rule Authoring

Similar skills