Case Study · SIG AI

I stopped trusting
AI benchmarks.
So I built my own.

SIG AI is an agent runtime I architected and built with AI in the loop. It runs open models on my own hardware, ranks every one of them against a test suite of real work, and refuses to let any of them touch my credentials. 4,390 lines of Python. No framework underneath it.

📊
87%
on a 15-task suite of real work, checked in code. Small n, so treat it as promising, not settled
🛡️
20 Bypasses
found by an adversarial audit aimed at my own permission gate, all closed
💻
100% Local
runs on a Mac mini, cloud tier disabled by default
sig — audit run
$ sig "audit this repo for exposed secrets"
SIG · LFM2.5-8B-A1B · llama.cpp :8086
tools: 7   mode: read-only   workspace: ~/projects/thebotique
→ grep "API_KEY|SECRET|TOKEN" ······· 12 matches
→ read .env.example ··················· ok
→ read .env ··························· DENIED
↳ secret shield: unreadable at the gate
→ read ~/.ssh/id_rsa ·················· DENIED
→ grep "token" --include=*.ts ····· 8 matches
3 findings · 13 tool calls · 4 denied by the gate · 100% tool validity

A real audit run. The agent tried to read credentials and the gate stopped it, four times.

The Problem

Everyone claims their AI setup works. Almost nobody can show you the number.

I had spent a year running a large agent system on a third-party framework. Dozens of agents, dozens of scheduled jobs, everything routed through cloud APIs. It was broad and it was genuinely useful. But I could not answer a basic question about it: which parts actually worked, and how well?

The industry answer is a leaderboard. Leaderboards measure models on academic tasks, not on whether a model can find the right file in my repo and change one line of it without breaking something else. That is the work. Nobody was measuring the work.

So I built the thing I actually wanted: a small runtime I own end to end, running open models on hardware sitting in my office, with a test suite made of real tasks and a permission system strict enough that I would let it near a production repo. Then I measured everything and let the numbers pick the defaults.

Worth saying plainly, because the line count invites the wrong assumption: I did not hand-type 4,390 lines of Python. I designed the architecture, made every call in it, and directed AI to write the code, then reviewed, tested and attacked the result until it held. Doing that well is the job I am describing on this site. This is the artifact that shows I can do it.

If you can't measure your AI system, you don't have a system. You have a vibe.

The Measurement

Four models. Five configurations. Fifteen real tasks.

The suite is 15 tasks drawn from work I actually do: read a file and explain it, count matches across a codebase, make an exact edit, find files by pattern, change something across multiple files. Every result is checked by an assertion written in code. No model grades its own homework.

LFM2.5 8B
llama.cpp · Q8_0
87%13/15
100% valid59 tok/s
Qwen3 8B
Ollama · default
60%9/15
100% valid21 tok/s
GPT-OSS 20B
LM Studio · default
60%9/15
98% valid10 tok/s
LFM2.5 8B
mlx_lm · 8-bit
40%6/15
96% valid48 tok/s
Gemma (local build)
llama.cpp · Q4_K_M
33%5/15
98% valid11 tok/s

The Finding

Rows one and four are the same model.

Identical weights. Identical quantisation. Identical tasks. On the same build, one runtime scored 40% and the other 67%, a 27-point spread that had nothing to do with the model at all. I ruled out the sampler configuration and it held. (The 87% in the table came later, after a separate bug fix, so it is not the engine comparison.) The entire industry benchmarks models. Hardly anyone benchmarks the software running them, and in my testing that mattered more than which model I picked.

The Architecture

Nineteen modules. No framework underneath.

SIG is a CLI, an interactive shell, and an HTTP daemon over the same core loop: messages in, model call, tool calls, results back. Six decisions define the whole thing.

🔌

One Protocol, No Adapters

Every backend is spoken to as OpenAI-compatible HTTP. The harness never loads a model itself. Swapping engines or porting the whole system to Linux is an edit to a YAML file, not a code change.

🪜

A Ladder, Not a Router

Six tiers configured, five in the routing ladder, ordered by measured performance. The tier order is the routing policy. When a task fails, the agent walks down the ladder. No heuristics, no guessing which model is good at what.

🔒

Three-Layer Permission Gate

Path scope confines writes to the workspace. A secret shield makes credentials unreadable at the gate. A command gate refuses destructive shell patterns. Denials come back as tool results, so the model can adapt instead of crashing.

⚛️

Kernel-Enforced Sandbox

Shell commands run under sandbox-exec on macOS and bubblewrap on Linux. The kernel enforces the boundary, not a regex. If the sandbox is unavailable it degrades loudly, never silently.

🧠

Two-Tier Memory

A 1K-token identity core loads on every session. The remaining 127 notes are searched on demand. Vault writes are off by default and confined to three folders. Years of personal data is not a write target for an 8B model.

📓

Everything Is Logged

Every gate decision, allow and deny, is written to an append-only audit log stamped with the session it came from. 3,384 records so far. If the system did something, there is a record of why it was permitted.

The Audit

I attacked my own system and it lost.

Writing a permission gate is easy. Proving it holds is not. So I pointed a script at mine whose only job was to break it, and it did, immediately and repeatedly. Every hole it found is documented, including one hypothesis of mine that the results flatly refuted.

01

An adversarial audit was written specifically to break the permission gate

02

It found 20 bypasses, including reading live API keys and writing into a protected archive

03

The root cause was structural: every path protection lived in the read and write checks, and shell commands never consulted them

04

All 20 were closed by moving enforcement below the tool layer, into the kernel sandbox

05

54 out of 54 gate tests now pass, and 15 out of 15 ordinary commands still run, so the fix did not over-block

06

Verified in the field: the security agent audited a production repo in 13 tool calls and the gate denied it 4 times

A safety layer you have never attacked is not a safety layer. It's a hope.

Design Decisions

Four things the measurements taught me.

The Engine Was Worth More Than the Model

The single most surprising result of the whole build. Identical LFM2.5 weights at identical quantisation scored 40% under one runtime and 67% under another on the same build. Same model, same quant, same tasks, a 27-point spread. I ruled out the sampler configuration and it held. Everyone benchmarks models. Almost nobody benchmarks the thing running them.

Capability Breadth Costs Accuracy

I added an eighth tool for line-level editing. The eval dropped from 10/15 to 7/15 across two consecutive runs, and the model only reached for the new tool 4 times against the existing one's 58. It cost selection accuracy on tasks that had nothing to do with editing. So the default tool block is deliberately lean, and everything else is one keystroke away instead of preloaded.

The Cloud Tier Is Commented Out on Purpose

A frontier cloud model is wired in and disabled by default. An always-available paid tier quietly becomes the tier that does everything, and then you have not built a local system. You have built an expensive proxy with extra steps. Turning it on has to be a decision, not a fallback.

Programmatic Checks, No Self-Grading

15 real-work tasks with assertions written in code. No model grades its own output. Run-to-run variance is roughly plus or minus two tasks, and the docs refuse to read anything smaller than that as signal. A benchmark you can talk yourself into is not a benchmark.

In Practice

What I actually use it for.

Voice notes that execute

A folder on my desktop is watched. Drop a voice memo in it and SIG transcribes it locally with Whisper and acts on what I said. Nothing leaves the machine and it costs nothing per use.

Security sweeps

The read-only sentinel agent audits a repo for exposed secrets and weak permissions. It reports findings and is structurally incapable of "fixing" them behind my back.

Screenshots as input

A vision model reads whatever is on my screen. Seeing and acting are two different models, so vision sits deliberately outside the routing ladder rather than in it.

Scoped repo work

Reading, grepping, explaining, and small exact edits, the read-mostly work the 87% number actually covers. Multi-file engineering still gets supervised. The README says so out loud.

The Takeaway

The point was never the agent. It was knowing whether it worked.

SIG AI is smaller than the system it replaced and does less on purpose. What it has instead is a number attached to every claim: which model, on which runtime, passing which tasks, at what speed, behind a gate that has been attacked and held. That shift, from breadth I assumed to depth I measured, is the whole project. The same instinct applies to a brand platform or a product roadmap. Build the thing, then build the test that tells you the truth about it.

Anyone can wire up an agent. The hard part is being able to prove what it does, and what it will refuse to do.

Let's talk →