Order Samurai
Order Samurai
PillarsRonin ModeProofPricing 389+ tests Open free dashboard
Govern the agents that work while you sleep

The missing discipline for agent fleets.

Order Samurai intercepts prompt injections, scrubs leaking credentials, and kills runaway spend across your coding-agent fleet — entirely on your machine, fail-closed by default, zero cloud telemetry.

Get started See the dashboard →
GOVERNS CLAUDE CODECODEX CLIGEMINI CLICURSORLANGGRAPHOPENCODE
[SECURITY]Prevented API token leakage in subagent spawn (Claude Code)
[SPEND]Halted runaway retry loop on Gemini Pro ($3,940 saved)
[AUDIT]Neutralized prompt injection attack vector in Cursor session
[RONIN]Auto-remediated 14 ATT&CK kill chain vulnerabilities
[VERIFIER]Enforced fail-closed sandbox policies across 42.5 agent hours
[SECURITY]Prevented API token leakage in subagent spawn (Claude Code)
[SPEND]Halted runaway retry loop on Gemini Pro ($3,940 saved)
[DOJO]Nightly Dojo completed 12 meditation cycles with 100% pass rate
[CRAFT]18.5 human review hours saved via automated doc-parity verifiers
[FLEET]Governing 5 model platforms: Claude, Codex, Antigravity, Cursor, Ollama
[SECURITY]Zero credentials leaked across 62 active agent sessions
[POLICY]Root hygiene policy & zero-trust boundaries fully enforced
[DOJO]Nightly Dojo completed 12 meditation cycles with 100% pass rate
[CRAFT]18.5 human review hours saved via automated doc-parity verifiers

Observability watches. Nobody fixes.

A stuck loop burns $6,000 overnight. A pasted README walks a credential out the door. Your tracing vendor records all of it beautifully — and bills you per trace for the footage. The fixing is still yours, at 2am.

Order Samurai closes the loop the market left empty: fail-closed hooks block the incident live, spend-capped reflexes remediate routine failures overnight, and an adversarial verifier signs every receipt.

01 · 剣 SWORD

Kill chains, cut mid-swing.

Fourteen ATT&CK-style chains monitored. Prompt injections intercepted in the moment; credentials and IP scrubbed before any commit carries them out. Fail-closed, not best-effort.

feat/agent-7 ✗ blocked
chain-13 prompt-injection via pasted README
chain-14 secret-scrub hit on ANTHROPIC_API_KEY
chain-02 sandboxed execution verified
ronin-daemon active · 23 tasks
spend-limit $40.00 / $50.00 daily cap
dojo-cycles 12 meditation runs · 0 regressions
agent-hours 42.5h returned to team
02 · 弓 BOW

Operations that run overnight, safely.

Nightly Dojo cycles run while you sleep. Task backlogs drain autonomously, spend caps hold hard at runtime, and failed steps auto-quarantine without breaking master.

03 · 筆 BRUSH

Spend caps that are enforced, not charted.

Hard budget ceilings, model-tier routing, token-execution density. The silent bleed stops where the budget says it stops — no dashboard required to notice.

projected burn by 07:00$3,940
fleet budget ceiling$120
spend cap engaged · agent #7 suspended
→ $3,820 never left the building.
CRAFT.md read on every review
slop density0.8% ▾
rework loops / PR1.2 ▾
doc parity94%
output that stays worth reviewing
04 · 芸 ARTS

Craft metrics for the human in the loop.

Slop density, rework loops, documentation parity. One score per discipline, never averaged — a failure in one pillar can't hide behind the other three.

AUTONOMOUS GOVERNANCE ARCHITECTURE

Ronin Mode & The Self-Improving Loop.

Observability without reflexes is just an expensive audit log. Order Samurai pairs continuous background Ronins with real-time reflex interception and overnight Dojo cycles — transforming reactive agent oversight into an autonomous self-healing engine.

01 · THE RONIN LOOP

Continuous Fleet Guardian

Operates as an autonomous background guardian across IDE namespaces and agent runtimes. Evaluates every tool execution and file change against the four-pillar governance contract without requiring active prompt engineering.

02 · REFLEX ALERTS

Defending the Floor

Deterministic, zero-latency interception. When a pillar metric degrades (prompt injection attempt, secret in staged diff, runaway spend spike), reflexes fire instantly to block the action and isolate the process before failures compound.

03 · DOJO CYCLES

Raising the Ceiling

Structured, autonomous training runs (Keiko) that run while you sleep. The Dojo processes task backlogs, executes automated regression sweeps, stages patches via maker-checker verification, and continuously recalibrates baselines.

SELF-IMPROVING FEEDBACK LOOP

From Traps to Calibrated Baselines.

Every intercepted prompt injection and completed Dojo work-unit feeds back into local calibration coefficients. The system learns the exact baseline performance of your fleet across Claude, Codex, Antigravity, and Cursor — making defenses sharper every night.

AGENTIC LOOP FLOW
1. REACTION → Intercept injection / scrub secret
2. ISOLATION → Stage patch to pending_remediation
3. DOJO CYCLES → Execute overnight keiko backlog run
4. CALIBRATION → Update empirical metrics & thresholds

Proof, labeled honestly.

Anonymized 30-day pilot fleet. A figure stays SIMULATED until twenty empirical samples exist — we'd rather print a gap than fake a graph.

$3,940
median runaway-spend incident intercepted before execution
● MEASURED
42.5h
agent time returned per week through overnight remediation
● MEASURED
212
secrets scrubbed pre-commit across 14 monitored kill chains
● MEASURED
42s
mean time to heal a routine degradation, unaided
○ CALIBRATING
47 metrics in catalog22 live reducers0 simulated numbers shown as live389+ passing tests The Honesty Invariant · audit it in the source

A flat price, kept flat.

No trace metering, ever. Your agents' verbosity is our problem to eliminate — not our revenue.

OSS CORE
$0 forever
Apache 2.0 open core. Full four-pillar scoring, fail-closed security hooks, and alert-only spend caps.
Try the free dashboard
PRO VERSION
$199 one-time
Active spend enforcement, Nightly Dojo, autonomous reflex remediation, maker-checker patch staging, and offline perpetual key.
Get Pro Lifetime ($199)

Get started

$curl -fsSL get.ordersamurai.dev | bash

Reads the Claude Code logs you already have and hands you a governance report in ten minutes — before any daemon runs. Requires nothing but the logs on disk.

Why one command, many harnesses
samurai Claude Code hooks + verifier
Codex CLI + codex-tuned kill chains
Gemini CLI + gemini-tuned kill chains
LangGraph daemon sidecar
Same governor, recompiled per target. Builds for harnesses with known failure modes carry extra reflexes for those habits.