RULEST v3.0 · MODULAR GPU‑ACCELERATED RULE EXTRACTION

Rulest

Extract high‑probability Hashcat rules by analyzing transformation patterns between base and target wordlists. Four‑phase GPU extraction, built‑in seed families (A–M), Token‑Strip pre‑pass (Phase 0), Genetic Algorithm (Phase 3), functional minimization, and a new CELF‑based greedy coverage selector — now shipped as a clean, testable Python package.

OpenCL GPU Four‑Phase ruleflow Seed Families A–M Functional Minimization Bloom Filter VRAM 5,600+ Validated Ops Token‑Strip Phase 0 Genetic Algorithm Phase 3 Modular Package CELF Greedy Selection
Overview

Rulest replaces naive BFS chaining with a GPU‑first four‑phase extraction engine. Given a base wordlist (source) and a target wordlist (dictionary), it reverse‑engineers the Hashcat rules that transform base words into target words — using OpenCL parallelism, a VRAM Bloom filter, four distinct extraction phases, and an optional Genetic Algorithm for deep‑chain discovery.

The result is a production‑ready .rule file, minimized via signature‑based functional deduplication. 100% compatible with Hashcat’s GPU engine (max 31 ops, no rejection rules). An optional Phase 0 Token‑Strip CPU pre‑pass reverse‑engineers exact chains from target passwords using 14 extraction modes before any GPU work begins.

v3 keeps every Stage 0–3 behavior identical to v2 in its default mode, but splits the old 210 kB single‑file engine into a proper rulest/ package and adds an optional post‑processing selector: Coverage Evaluation + Greedy CELF Selection, which orders and trims rules by real, verified marginal password recovery instead of raw GPU hit count.

Core Features
⚡
Four‑phase extraction ruleflow
Phase 0 Token‑Strip (optional CPU) → Phase 1 (single‑rule sweep) → Phase S (built‑in seed families) → Phase 2 (hot‑rule biased chains) → Phase 3 GA (optional).
🌱
Built‑in seed families (A–M)
Thirteen deterministic families: digits, date patterns, special chars, leet substitutions (J), double‑transform chains (K), special‑before‑digit (L), leet+transform combos (M).
🧠
Bloom filter GPU lookup
FNV‑1a hash‑based membership test in VRAM (16–256 MB) for instant target validation.
🔬
Functional minimization
Signature‑based deduplication using probe words — removes 20–60% of equivalent rules without losing coverage.
🗜
Phase 0 — Token‑Strip (14 modes)
Optional CPU pre‑pass: reverse‑engineers exact rule chains from target passwords using 14 extraction modes with multiprocessing. Singles → Phase 1; chains → Phase S + Phase 2.
🧬
Phase 3 — Genetic Algorithm
Optional evolutionary search (--genetic) with 2× novelty bonus for new chains, 20% time reservation, stagnation guard, and tournament selection + crossover + mutation.
📊
Hit counting & ranking
Per‑rule hit frequency from GPU validation; default output sorted by frequency (v2 behaviour).
🎯
New — Coverage Evaluation + Greedy CELF Selection
Each candidate rule is verified against the real target wordlist to build a rule → distinct‑recoveries map, then rules are chosen one‑by‑one by largest marginal gain (lazy greedy / CELF) to maximise union coverage per rule budget.
Extraction ruleflow (v3)
Phase 0
Token‑Strip pre‑pass (optional)
CPU‑only pre‑pass (--token-strip) with 14 extraction modes and multiprocessing. Reverse‑engineers exact rule chains from target passwords; singles → Phase 1; chains → Phase S + Phase 2.
Phase 1
Single‑rule sweep
5,600+ validated single rules applied to every base word; Bloom filter checks membership in target. Collects effective single rules and hit counts.
Phase S
Built‑in seed families (A–M)
Thirteen predefined families covering digits, date patterns, special chars, leet substitutions (J), double‑transform chains (K), special‑before‑digit (L), and leet+transform combos (M). Depth 2–9 tested as dedicated GPU pass.
Phase 2
Hot‑biased chain generation
Chain builder with 60% hot‑rule bias (using Phase 1 winners), 30% seed extension, 10% random exploration. Respects per‑depth budgets and max 31 ops.
Phase 3
Genetic Algorithm (optional)
Evolutionary search (--genetic) with 2× novelty bonus, dedicated 20% time reservation (min 120s), stagnation guard, tournament selection + crossover + mutation.
Post‑processing
Minimization & export
Probe words compute functional signatures; SQLite‑backed dedup for >500k candidates; keep highest‑hit rule per equivalence class.
Selection
Rule selection (rulest/selection.py, new in v3)
--select-mode frequency (default, identical to v2) or --select-mode greedy for CELF marginal‑coverage selection against the real target, with optional --select-budget/--select-budgets and --select-cost-alpha.
Built‑in Seed Families (Phase S) — A through M
A – Pure Prepend
Digits prepended (^digit) depths 1–4
B – Pure Append
Digits appended ($digit) depths 1–4
C – Mixed Prepend/Append
All combos of ^d and $d depths 1–4
D – Transform + Digit/Bracket
l,u,c,C,t,… + 1‑4 digits or [ ]
E – Date Patterns
DDMM, MMDD, YYYY, DDMMYY, DDMMYYYY + transforms (up to depth 9)
F – Append Special Chars
Top‑15 specials appended ($char) depths 1‑3
G – Prepend Special Chars
Top‑15 specials prepended (^char) depths 1‑3
H – Transform + Special Char
Transform + 1‑2 specials (append/prepend), depths 2‑3
I – Digit(s) + Special Char
Digits + core specials (!@#$%*?) depths 2–4
J – Leet Substitutions
Top‑10 leet pairs (sa@, se3, so0, si1, sl1, ss5, ss$, st7, sa4, si!) — pure + digit/special combos
K – Double‑Transform Chains
All 225 ordered pairs of structural transforms (l, u, c, C, t, r, d, f, E, k, K, {, }, [, ])
L – Special‑before‑Digit
Special char first, then digits — covers word!12 and !12word patterns (depths 2‑3)
M – Leet + Transform
Every leet op paired with every structural transform in both orderings (≈300 chains)

Disabled with --no-builtin-seeds. These seeds run as a dedicated phase and are also forwarded to Phase 2 as scaffolding for deeper chains.

Modular Package Architecture (New in v3)

The original 210 kB single‑file rulest_v2.py engine has been split into a clean, testable Python package. Behavior of Stages 0–3 and frequency‑based output remains identical to v2 in the default selection mode.

rulest/state.py
Runtime flags
rulest/common.py
Logging, colors, constants, rule validator, CPU applicator
rulest/minimize.py
Signature‑based functional minimization
rulest/token_strip.py
Phase 0 (CPU exact‑match)
rulest/devices.py
OpenCL device discovery & dynamic sizing
rulest/kernel_source.py
OpenCL C kernel template
rulest/gpu_engine.py
Single‑GPU engine
rulest/gpu_worker.py
Process‑isolated GPU worker
rulest/multi_gpu.py
Multi‑GPU orchestration
rulest/genetic.py
Phase 3 genetic algorithm
rulest/extractor.py
Pipeline orchestrator
rulest/selection.py
New — coverage evaluation + CELF selection
rulest/cli.py
Argument parsing & main flow
run_rulest.py
Entry point
Toolchain Integration & Workflow

Rulest v3 integrates seamlessly with the A1131 ecosystem:

Rulest (v3)
Produces a minimized .rule file, ordered by frequency (default) or by CELF greedy coverage selection.
Minimizer (optional)
For cross‑corpus signature minimization across multiple rule sets.
Ranker / Aether
Re‑benchmark or visualize rule performance on different validation sets.
💡 Tip: Rulest already removes functionally equivalent rules via signature minimization (21+ probe words). To further deduplicate across multiple runs, pipe the output through Minimizer with a unified probe set. For the most coverage‑efficient ruleset, run with --select-mode greedy instead of the default frequency ordering.
# Typical workflow (v2 behaviour, default)
python3 run_rulest.py rockyou.txt targets.txt --max-depth 3 -o extracted.rule
# Optional cross‑minimization
python minimizer.py --probe-words probes.txt extracted.rule final.rule
⚙️ GPU & Performance
  • VRAM‑aware batch sizing (baseline 8 GB, scales down to 4 GB)
  • Dynamic Bloom filter: 16–256 MB based on target size
  • Multi‑device support: --list-devices, --device index/name
  • Per‑depth chain budgets (--depth2-chains … --depth10-chains)
  • Max rule ops: 31 (Hashcat GPU limit), rejection rules automatically excluded
📦 Requirements & Installation
Python ≥3.8 · numpy · pyopencl · tqdm
# Clone & install
git clone https://github.com/A113L/rulest.git
pip install -r requirements.txt
python3 run_rulest.py --help

OpenCL 1.2+ GPU (NVIDIA, AMD, Intel). CPU fallback supported but slow. Unit tests included under tests/ — install requirements-dev.txt and run pytest.

Quick Examples
View rulest on GitHub
Basic extraction — depth 2
python3 run_rulest.py rockyou.txt target.txt --max-depth 2 -o myrules.rule
Specific GPU device + custom chain budget + time limit
python3 run_rulest.py base.txt dict.txt --device 0 --depth3-chains 80000 --target-hours 1.5
External seed rules, skip built-in families, depth 4
python3 run_rulest.py base.txt target.txt --seed-rules previous_seeds.txt --no-builtin-seeds --max-depth 4
Phase 0 — Token-Strip pre-pass (14 modes, multiprocessing)
python3 run_rulest.py rockyou.txt target.txt --token-strip --max-depth 4 --target-hours 1.0 -o with_ts.rule
Phase 3 — Genetic Algorithm (novelty-weighted, 20% time reservation)
python3 run_rulest.py base.txt target.txt --max-depth 3 --target-hours 1.5 --genetic -o evolved.rule
New in v3 — Greedy CELF selection, keep best 50k rules
python3 run_rulest.py base.txt target.txt -o rules.txt --select-mode greedy --select-budget 50000
New in v3 — Multiple budget sizes + exact recovery report in one run
python3 run_rulest.py base.txt target.txt -o rules.txt --select-mode greedy --select-budgets 64,250,1500,10000,25000,50000,150000 --exact-recovery

Output includes header with total candidates, minimization stats, and per‑depth rule counts. Sorted by GPU hit frequency (default) or by CELF marginal coverage (--select-mode greedy).

New CLI Parameters — Rule Selection (Post‑processing)
--select-mode {frequency,greedy}
frequency (default): sort by raw GPU hit count, identical to v2 behaviour. greedy: CELF marginal‑coverage selection against the real target.
--select-budget N
Max rules to keep when using greedy (0 = run to saturation).
--select-budgets LIST
Comma‑separated budgets, e.g. 64,250,1500,10000,25000,50000,150000. Emits extra files <output>.<budget>.txt cut from the same ordering.
--select-cost-alpha A
gain(r) = new_recovery / depth(r)**A. 0 = pure marginal coverage (default); >0 favours shorter/cheaper chains.
--exact-recovery
Report exact distinct target‑word recovery after greedy selection.
Get the Source

Rulest is open source and hosted on GitHub. MIT licensed.

View rulest on GitHub
Summary & Upgrade Notes — v2 → v3

v3 is a drop‑in upgrade for all users: same GPU engine, same Stages 0–3, same default output — now organized as a package, plus an optional smarter rule selector for anyone who wants tighter, more diverse rulesets.

What Changed
Structure
Single file → modular rulest/ package
The 210 kB rulest_v2.py is now 15 focused modules plus a run_rulest.py entry point, with unit tests under tests/.
New
Coverage evaluation + CELF greedy selection
v2 sorted purely by raw GPU hit count, which often wasted budget on redundant rules recovering the same passwords. v3 can verify each rule's real, distinct target recoveries and greedily maximise coverage per rule.
New
rulest/selection.py
Powers --select-mode, --select-budget(s), --select-cost-alpha, and --exact-recovery.
Compatibility
All users
Default mode unchanged
--select-mode frequency (default) produces the same ordering and output format as rulest_v2.py.
All users
Every previous CLI flag preserved
Multi‑GPU, genetic algorithm, token‑strip, and bloom filter options work exactly as in v2.
Optional
Greedy selection is opt‑in
Enable with --select-mode greedy when you want CELF marginal‑coverage ordering against the real target instead of raw frequency.
💡 Note: Entry point is now run_rulest.py instead of rulest_v2.py. All existing flags (--max-depth, --target-hours, --bloom-mb, etc.) behave identically to v2.