Skip to main content
Projects

Vonnis

Offline script triage, measured against real corpora

A script triage engine that reads a suspicious PowerShell or shell script and answers, in plain English, whether running it is a bad idea. 111 deterministic YAML detection rules, no API key and no network for the verdict, and a measurement harness that scores the rules against 2,526 real package maintainer scripts and 1,862 Atomic Red Team tests.

Status
Active
Touches
Python / Detection Engineering / MITRE ATT&CK / YAML Rule Engine / asyncio / SQLite / Docker
References

What it does

Vonnis reads a suspicious script and tells you whether running it is a bad idea, in a sentence a non-technical person can act on. Point it at a .ps1 attachment or paste a shell one-liner, and you get a verdict instead of a wall of technical output.

It is not built for security researchers, who already have Ghidra and do not need it. It is built for the person on the other end: the IT tech who just received a PowerShell file from an unknown sender, or the business owner whose employee forwarded something that looked off. If you do not know what ExecutionPolicy Bypass means, that is the point.

The verdict, the rules that fired, the step-by-step walkthrough and the extracted indicators are all produced on your machine. No API key, no account, no network call.

  • Interfaces: command line, Telegram bot, Docker
  • Detection: 111 YAML rules across ten ATT&CK-aligned categories
  • Walkthrough: 67 neutral action descriptions, also YAML data
  • Suite: 527 tests, running offline with no keys
  • Measured: 0 false positives over 2,526 real benign scripts

How the verdict is decided

Two layers, and the order is the design.

The rules decide. Deterministic pattern matching kept in YAML files sets a risk floor. Reputation can only raise it. A VirusTotal or MalwareBazaar hit sets a second floor, tiered by how many vendors actually agreed, because one engine out of seventy is a routine false positive and sixty-two is not. The floors combine by taking the higher, never by adding, so a URL that both trips a rule and is flagged externally is one fact confirmed twice rather than two pieces of evidence.

A model, if you configure one, only rewords. It can raise the risk and write a better summary sentence. It cannot lower the verdict and it does not write the walkthrough. A script controls its own contents, so it is framed to the model as untrusted data inside a per-call random fence the script cannot forge, and anything in it addressed to the model is reported rather than obeyed.

Rules do not score independently. Each one maps to a behaviour — fetching remote code, hiding a command, keeping itself running, reaching for passwords — and the classifier scores behaviours, capped at two hits each. Five overlapping regexes for the same download cradle cannot stack into a critical verdict on their own, and a single high-severity signal with nothing corroborating it comes back as “investigate further” rather than “do not run”.

That cap is what separates the official Deno installer from a dropper. Both are curl … | sh.

Measuring the rules instead of trusting them

The repository ships a sixty-sample corpus, and it is deliberately described as a regression guard rather than a measurement. Sixty samples cannot support an accuracy claim; they exist so a rule change cannot quietly break something that used to work.

The real instrument is scripts/evaluate.py, which scores the rules against corpora nobody wrote for the purpose: every package maintainer script on the machine, and Atomic Red Team.

Because every rule carries the ATT&CK techniques it claims and every atomic carries the technique it emulates, the harness can report the number that actually matters: which techniques a rule claims to cover but never detects. Those are the coverage holes, stated by the tool against itself.

It works. CRED-WALLET-001 matched the bare word keystore and flagged the stock ca-certificates-java package as CRITICAL. Forty hand-written samples never hit it. Two thousand real ones did, on the first run.

The last round of work was driven entirely by that harness — 26 new rules, none written from imagination, each closing a gap the measurement surfaced:

beforeafter
Detection rules85111
False positives on benign corpus10
Atomics with a rule firing245 (13.2%)411 (22.1%)
Serious silent misses11172
Tests505527

Fix the harness before you trust it

The most useful thing that round taught me had nothing to do with rules.

The “claims a technique but never detects it” report was rolling sub-techniques up to their parent, so a rule claiming T1222.002 looked responsible for the Windows takeown atomics it had never claimed. Three of the six reported failures were the instrument, not the rules. An evaluation harness is code, it has bugs like any other code, and a measurement you have not audited will send you off fixing things that were never broken.

The three genuine false positives were all one bug wearing different clothes: a tool name matching inside a longer name. keystore inside ca-certificates-java. bcrypt inside python3-bcrypt.postinst, which is password hashing rather than file encryption and did not belong in the list at all. An offensive-tooling rule hitting eight Kali package scripts byte-compiling the tools they ship, because installing a tool is not running it.

Each fix carries a test holding both halves — the true positive and the benign lookalike — because the suppressors are the part that breaks quietly.

Saying “I don’t know”

A script Vonnis could not parse produced exactly the same empty result as a script with nothing wrong in it. Both came back NO PROBLEMS FOUND with exit code 0. A Go dropper now returns COULD NOT DETERMINE and exit 1 instead.

The threshold for that came from measurement, and the obvious version of it was wrong. “Recognised nothing” on its own is not a signal of blindness: it is true of 46% of the package maintainer scripts on a normal machine, whose clean verdicts are correct. Gating on it would have flipped 40% of them to no-answer while catching none of the Atomic Red Team misses, which are all one-liners. Three conditions together are the signal — the language was not identified, no rule matched, and nothing the script does was recognised — and that covers 1.3% of real benign scripts while catching the case the check exists for.

The false positives that are reported anyway

Ten of the forty benign samples land on MEDIUM / INVESTIGATE FURTHER, and that includes the genuine Chocolatey and Deno installers. To the IT tech this is built for, an “investigate further” on the official Deno installer is a false alarm, so the README reports it rather than rounding it down to zero.

They land there because they really do pipe a downloaded script straight into a shell. Suppressing that to make the table look better would be the wrong trade. The benign half of a corpus is the half that matters: a false positive on the real Chocolatey, Rust, Homebrew, nvm, pyenv or Docker installer teaches people to ignore every future verdict.

Design decisions worth naming

No vendor SDK for the model. The client speaks the OpenAI chat-completions dialect over plain HTTP, so one adapter covers Groq, Ollama, vLLM, OpenRouter and any in-house gateway. An organisation that cannot send scripts to a third party points it at their own network. Being natively async also fixes the timeout properly; the previous client parked a blocking SDK in the default executor, where a timeout would have leaked the thread rather than stopped the work.

Rules and walkthrough steps are separate vocabularies. An action says what happened and is neutral — “downloads a file from the internet” is equally true of the Rust installer and of a dropper. A rule says whether to worry. A test fails the build if an action’s sentence presumes the answer, because mixing the two is how a walkthrough starts reading like an accusation.

Rules are data, not code. Contributing a detection means writing YAML with patterns, co-signal requires:, suppressor absent: clauses and the ATT&CK ids it claims — no Python. --rules-dir points the engine at an entirely different pack.

Addresses are defanged in every report, so nothing written about a malicious URL is itself a way to reach it.

Exit codes mean something — 0 nothing found, 1 suspicious, 2 malicious, 3 could not analyse — so it composes with other tools.

Stack

Python 3.10+, with a deliberately small core: aiosqlite and PyYAML are the only required dependencies, which is what lets vonnis analyze --offline run on a minimal install. Enrichment, the model client and the Telegram bot are optional extras.

Around 5,800 lines of Python and 2,600 lines of YAML data. ruff and mypy in CI, with disallow_untyped_defs on vonnis/core since that is the part every interface depends on. Ships with a Dockerfile and a compose file for the bot.