GPU-Accelerated Rule Ranking & Coverage Selection

Ranker v5.2

Now a full package: Multi-Armed Bandit ranking, fast CSV analysis, and CELF greedy max-coverage post-processing behind one dispatcher CLI.

OpenCL GPU MAB Algorithm Memory-Mapped I/O NumPy Optimized CELF Coverage pip-installable
3Tools, 1 CLI
50×Loading Speed
3Coverage Strategies
5v5.2 Bug Fixes
Overview

Ranker answers the question: "Which of my rules actually cracks the most hashes — and which minimal subset of them covers everything cracked?" — without running a full Hashcat session for every rule.

The package now wraps three previously-standalone scripts into one installable Python package with a single dispatcher CLI, run_ranker.py. rank allocates GPU compute via a Multi-Armed Bandit, spending more trials on promising rules and eliminating poor performers early. handler summarizes ranking-output CSVs. postprocess runs a GPU/CPU CELF greedy max-coverage selection over the top candidates to find the smallest ruleset that keeps the same crack coverage. None of the original logic was rewritten — each script keeps its own full argument list and remains runnable directly as a module.

Core Features
∞
Multi-Armed Bandit + Early Elimination
Allocates GPU compute to top performers. Low-performing rules eliminated after confidence threshold crossed.
⚡
Memory-mapped file loading
OS-level mmap for near-instant rule file loading. 50× faster than standard file I/O for large rule sets.
∫
NumPy-optimized scoring
Vectorized hit counting and confidence interval math using NumPy broadcasting.
⊞
All Hashcat modes
Compatible with hashcat rules attacks and custom corpora.
⚑
Rule validation on load
Rejects rules using banned operators (M 4 6 X < > ! / ( ) = % Q) at load time, matching rulest's behaviour.
▣
CELF max-coverage post-stage
postprocess picks the smallest rule subset from your top candidates that preserves total crack coverage.
Ranking Workflow
Step 1
mmap rule files
Map all rule files into virtual memory. Load time: milliseconds even for 16K rules.
Step 2
GPU batch initialization
Compile OpenCL kernel, allocate device buffers for rules and hash corpus.
Step 3
MAB exploration phase
Each rule gets initial trials; bandit algorithm tracks win rates with UCB confidence bounds.
Step 4
Early elimination
Rules whose upper confidence bound falls below threshold are dropped, freeing GPU cycles.
Step 5
Final exploitation
Remaining budget spent on top-k rules for precise hit rate estimates.
Step 6
Export sorted ruleset
Output .rule file sorted descending by hit rate, ready for handler or postprocess.
One Dispatcher, Three Tools

Everything after the subcommand passes straight through to that tool's own unchanged argparse parser, so existing flags and docs for each script still work exactly as before.

rank
GPU rule ranking (MAB or --legacy exhaustive v3.2). --list-devices to pick a GPU.
handler
Fast CSV/rule-file analysis and summarization of ranker output. No GPU dependency.
postprocess
GPU/CPU CELF greedy max-coverage selection over the top candidates from a ranking run.
python3 run_ranker.py rank -w wordlist.txt -r rules.rule -c cracked.txt -o out.csv
python3 run_ranker.py handler -i out.csv -o summary.txt -r clean.rule -t 1000
python3 run_ranker.py postprocess -r out.csv -w wordlist.txt -k cracked.txt -o celf_selected.rule --candidates 20000 --budget 5000

Each tool also remains runnable directly as a module (python3 -m rule_ranker.ranker ...), and a run-ranker console script is available after pip install -e ..

Coverage Strategies (postprocess --strategy)

Which strategy wins depends on how sparse your rules' real coverage is and how small --budget is relative to --candidates — benchmark on a representative slice before committing to one for a long run.

▦
bitmap (default)
Builds an (n_candidates × cracked_universe) bitmap matrix, streamed to a memmap with hybrid dense/sparse row packing. One-time GPU cost; can still be too big for RAM/disk on very large pools.
◇
sparse
Stores each candidate's coverage as an explicit array of the indices it hits (dict or SQLite past --sparse-disk-threshold, default 100,000). One GPU pass, then CELF runs entirely on CPU — a good default when bitmap uses too much memory.
↻
recompute-gpu
Never builds a coverage matrix; each greedy round rescores still-live candidates on the GPU with upper-bound pruning. Wins when --budget is much smaller than --candidates — otherwise can do more total GPU work than bitmap.

--bitmap-path, --keep-bitmap, --no-hybrid, --in-ram, --no-parallel-celf, --celf-workers, --celf-io-threads and --celf-batch-multiplier are all ignored under recompute-gpu/sparse — no bitmap/matrix stage ever runs for those.

Get the Source

Ranker is open source and hosted on GitHub. MIT licensed.

View ranker on GitHub