TL;DR: How to Write YARA Rules That Catch Ransomware Families
Writing reliable YARA rules for ransomware comes down to one disciplined loop: extract static indicators—unique strings, byte patterns, ransom note text, mutexes, and import signatures—from known samples, combine them with counting conditions like 4 of them and metadata guards, then iterate against a corpus of known-bad and known-good files using yara-python until your confusion matrix is clean. Rules that match everything are noise; rules that miss variants are theater. The lab below gives you both halves.
Lab Setup: Installing YARA and yara-python
YARA—the pattern matching engine built by Victor Alvarez of VirusTotal and now maintained under the VirusTotal YARA project—ships as a CLI binary and a Python binding. For this lab you want both: the CLI for quick ad-hoc scans, yara-python for scripted testing pipelines.
Install on Linux:
sudo apt-get install yara
pip install yara-python
python3 -c "import yara; print(yara.__version__)"
yara --version
Keep the two versions aligned—yara-python bundles its own libyara, and mismatched rule syntax support (e.g., pe module features) between the CLI and the Python binding is a classic source of silent failures.
Organize your lab like this:
samples/malicious/<family>/— known ransomware samples per familysamples/benign/— Windows binaries, installers, common utilitiessamples/holdout/<family>/— variants reserved for validation, never used during authoringrules/— your YARA rule files, one per familyscan.py— the test harness we build below
On sourcing samples: theZoo, MalwareBazaar (abuse.ch), and CTF archives are legitimate starting points. Handle live samples in isolated VMs—ransomware does not respect your file system’s feelings.
YARA Rule Anatomy Refresher: Strings, Conditions, and Modules
A YARA rule has three blocks:
rule example_family_ransomware
{
meta:
author = "analyst@hmmnm.com"
family = "example"
version = "1.0"
strings:
$note = "ALL YOUR FILES HAVE BEEN ENCRYPTED"
$ext = ".l0ck3d"
$hex = { 6A 40 68 00 30 00 00 6A 14 8D 91 }
$mutex = "GlobalExampleLocker" ascii wide
condition:
uint16(0) == 0x5A4D and 3 of them
}
Key mechanics mid-level analysts should already know but often underuse:
ascii widecatches both encodings;nocaseloosens case sensitivity.- Hex wildcards (
??) and jumps ([4-8]) tolerate compiler variation. uint16(0) == 0x5A4Dis the MZ header check—your first line of defense against scanning text files.- The
pemodule exposes imports, sections, and timestamps:pe.imports("kernel32.dll", "CryptEncrypt"). - The
mathmodule gives you section entropy:math.entropy(0, filesize) > 7.5.
Collecting Ransomware Indicators: What to Extract Before Writing
The good news about ransomware as a rule-writing target: it’s chatty. Unlike stealthy implants, ransomware wants to talk to its victim. Before touching the keyboard, pull these static indicators—no detonation required:
- Ransom note text — exact phrases, contact emails, Tor URLs. Note filenames like
RESTORE-FILES.htmlor!README.txt. - Encrypted file extension lists — families hardcode hundreds of target extensions; these blocks are unique and large.
- Mutex and pipe names — created to prevent double-encryption; usually static per family or variant.
- Unique strings — author handles, error messages, base64 blobs, wallet addresses, version markers.
- Import signatures — clusters like
CryptEncrypt,WriteFile,FindFirstFileW,DeviceIoControl(for shadow copy deletion via VSS). - Packer artifacts — section names, high entropy, tiny import tables.
CISA and the FBI’s joint advisories on major ransomware groups—including the StopRansomware series—publish exactly these indicators. Use them; they’ve already done the triage across dozens of samples.
Writing Your First Family-Specific Rule
Here’s a representative family rule, built from static indicators typical of a locker variant:
rule Ransomware_ExampleLocker_Strings
{
meta:
author = "hmmnm.com detection lab"
family = "ExampleLocker"
version = "1.1"
date = "2025-01"
description = "Static indicators for ExampleLocker locker variant"
reference = "internal lab analysis"
strings:
$note_html = "YOUR FILES HAVE BEEN ENCRYPTED - READ THIS" ascii wide
$note_txt = "Send 0.1 BTC to" nocase ascii wide
$ext_block = ".exe.docx.jpg.png.pdf.xlsx" ascii
$mutex = "Global\ExampleLockerMtx" ascii wide
$tor = "http://examplelocker" ascii
$vss = { 76 73 73 5F 64 65 6C 65 74 65 } // "vss_delete" stub
condition:
uint16(0) == 0x5A4D
and filesize < 5MB
and 3 of ($note_*)
and $mutex
and 2 of ($ext_block, $tor, $vss)
}
Line-by-line logic:
- MZ header check scopes the rule to PE files first—cheap, eliminates 95% of false-positive surface instantly.
- filesize guard keeps the rule from matching huge archives containing incidental strings.
3 of ($note_*)demands corroboration across ransom note variants; no single string triggers detection.$mutexas a hard requirement anchors the rule because mutex names are rarely reused accidentally.- The trailing
2 ofgroup adds secondary corroboration without over-constraining.
Using yara-python to Test Rules Against Sample Sets
This is where most rules die or survive. Script the test:
import os, yara
rules = yara.compile(filepath="rules/examplelocker.yar")
def scan_dir(path, label, results):
for root, _, files in os.walk(path):
for f in files:
fp = os.path.join(root, f)
try:
matches = rules.match(fp)
except yara.Error:
continue
for m in matches:
print(f"[{label}] {fp}: matched strings:")
for s in m.strings:
print(f" {s.identifier}: {s.instances[:1]}")
results.append((label, bool(matches)))
results = []
scan_dir("samples/malicious/examplelocker", "malicious", results)
scan_dir("samples/benign", "benign", results)
scan_dir("samples/holdout", "holdout", results)
tp = sum(1 for l,h in results if l=="malicious" and h)
fn = sum(1 for l,h in results if l=="malicious" and not h)
fp = sum(1 for l,h in results if l!="malicious" and h)
tn = sum(1 for l,h in results if l!="malicious" and not h)
print(f"TP={tp} FN={fn} FP={fp} TN={tn}")
Every false positive prints its matched strings, which tells you exactly which indicator is dirty. Every false negative tells you which sample needs a looser pattern or belongs to a different variant. That feedback loop is the entire discipline.
Reducing False Positives: Condition Tuning and Thresholds
False positives come from indicators that looked unique but aren’t. Attack them in order:
- Raise string counts:
4 of theminstead ofany of them. - Add
pemodule checks:pe.imports("advapi32.dll", "CryptAcquireContextW")filters out non-PE-adjacent noise immediately. - Keep filesize guards; a 3GB installer tripping on a two-byte pattern is pure scan-time waste.
- Test against a large benign corpus—stock Windows DLLs, common installers, and games are notorious for incidental string collisions. VirusTotal’s YARA blog posts document real-world overmatching cases worth studying.
The rule of thumb: no indicator is trustworthy alone. Require corroboration.
Avoiding False Negatives: Variants, Obfuscation, and Generality
False negatives mean your rule is decorative. Ransomware families iterate fast—new builds rotate mutex names, recompile with different compilers, or swap note text. Countermeasures:
- Loosen hex patterns with wildcards and jumps:
{ 8B 45 ?? 50 [2-6] FF 15 }survives register and call-site variation. - Use
nocaseandwidejudiciously—note text often moves between encodings between builds. - For packed samples, static strings are invisible. Detect the packer stub or the entropy profile instead, or accept the gap and document it.
When to go variant-specific vs. family-level: if the strings you extracted appear across all observed builds and threat intel reports, write the family rule. If the only stable indicators are mutex names and note text that churn per build, write narrow variant rules under a shared naming convention (Ransomware_Family_Variant_A) so operators can distinguish coverage from your rule names alone. A family rule with a 20% miss rate is worse than three variant rules at 100%—you can’t hunt what silently passes.
Performance and Hygiene: Peephole Optimization and Scoping
At scale—thousands of rules scanning endpoints—performance is a security requirement, not a nicety:
- Front-load cheap conditions:
uint16(0) == 0x5A4Dand filesize guards short-circuit before expensive string evaluation. - Avoid pathological regexes.
/(a+)+$/-style patterns cause catastrophic backtracking; YARA will refuse some, but others just crawl. - Prefer literal strings and hex patterns with jumps over regex where possible.
- Follow the community YARA-Rules style guide: consistent naming, populated meta block, one logical family per file.
- Never use private rules as a substitute for scoping—they still execute.
Validating Rules Like an Analyst: Peer Review and Multi-Corpus Testing
Your holdout set exists for one moment: final validation. After tuning against your working corpus, run the rule against held-out variants you never touched. If it misses them, your rule is overfit. Standard validation practice:
- Document the detection logic—what each string represents—in meta fields, so the next analyst can update it when the family rotates.
- Version rules and date them; stale rules are silent debt.
- Run yarGen to auto-generate candidate strings from your sample set and cross-check whether you missed better indicators.
- Use yaralyzer to visualize whether your patterns match what you think they match.
- Peer review before deployment—two analysts reading a condition catch the overbroad wildcard that your corpus didn’t.
Deploying Rules: from yara-python Scripts to EDR and Threat Intel Feeds
Once validated, the same rules deploy across your stack:
- Scheduled scanning pipelines — your yara-python script becomes a cron-driven scanner over file shares and quarantine directories.
- Retrohunts — VirusTotal VT Intelligence retrohunt runs your rule against months of historical samples to measure prior prevalence and catch missed incidents.
- EDR custom detections — most major EDR platforms ingest YARA for on-access or scheduled scans.
- File-event integration — osquery file events or msticpy-enriched pipelines can trigger YARA scans on writes to sensitive directories, pairing behavioral file-event telemetry with static matching.
Mature programs pair YARA with behavioral detection—YARA catches the file at rest, EDR catches the encryption burst in memory. CISA’s StopRansomware guidance emphasizes exactly this layered posture.
Complete Lab Checklist and Next Steps
- ☐ YARA CLI and yara-python installed, versions aligned
- ☐ Sample directories: malicious, benign, holdout
- ☐ Indicators extracted (notes, extensions, mutexes, imports, hex)
- ☐ Rule authored with MZ check, filesize guard, corroboration conditions
- ☐ scan.py run; FP and FN counts at zero on working corpus
- ☐ Validated on holdout variants
- ☐ Peer reviewed, versioned, meta documented
- ☐ Deployed to at least one pipeline (retrohunt or scheduled scan)
Next labs worth your time: apply the same loop to APT tooling (the indicators are subtler but the process is identical), then move to memory scanning—YARA against process memory dumps catches string-dyed and packed malware that file scanning misses. The author-test-tune loop doesn’t change; only the substrate does.
Frequently Asked Questions
Can YARA detect packed or encrypted ransomware?
Generally no—not on the encrypted payload itself. If the ransomware body is packed, its strings and imports are hidden from static analysis. You can still detect the packer stub, section entropy anomalies (via math.entropy), or mismatched section characteristics, and you can detect encrypted files at rest via extension and note-file patterns. But be honest in your rule’s meta: a rule that only catches the packer is not a family detection.
Is writing YARA rules for malware legal?
Yes—authoring, publishing, and sharing YARA rules is legal and is standard threat intelligence practice worldwide. The constraints sit elsewhere: obtaining and storing live malware samples may implicate local law, corporate policy, and safety obligations. Use controlled repositories like theZoo or MalwareBazaar, work in isolated environments, or use indicator-only workflows (writing rules from published IOC reports) to stay entirely clear of live sample handling.
What’s the difference between yara-python and the YARA CLI?
Same engine, different interface. The CLI is ideal for one-off scans and quick iteration; yara-python gives you programmatic control—compiled rule objects, match objects exposing matched strings and offsets, and clean integration into scanning pipelines, alert triage, and CI-style rule testing. If a rule will ever run without a human watching the terminal, write the harness in yara-python.
How many strings should a YARA rule require to match?
Rule of thumb: require several unique indicators, not one. Something like 4 of ($a*, $b*, $c*) across distinct indicator classes—note text, mutex, extension block—balances detection sensitivity against false positives. A single-string rule is a false-positive generator; a six-condition rule on a small variant is a false-negative machine. Corroborate, but don’t over-constrain.
How do I test YARA rules without real malware samples?
Three options: build synthetic files containing the extracted strings (a text file with your ransom note and extension list embedded will exercise the string logic, though not MZ/module conditions); use indicator-only workflows authored straight from CISA or vendor reports; or source CTF sample sets, which are explicitly distributed for safe practice. Always pair with a benign corpus—negative testing matters as much as positive.
Related reading
- Agentic AI Security Cheat Sheet: Threat Models, MCP Hardening Checklist and Detection Hooks
- Alert Fatigue Is a Design Problem: Building a Detection Engineering Lifecycle That Survives Contact
