Verify after an action
Checks the result actually satisfies what was asked.
A judgment layer between your coding agent and its tools. It pauses risky actions, blocks injected instructions, and trims tool output to what the task needs, before the main model ever sees it.
bun add @brainstem/reflexes @brainstem/pi-adapter
Scripted replay of a real session shape. No live API calls; token counts are illustrative. Switch to "without" to see the same session unguarded.
less tool output sent to the model
of raw-output accuracy in the Focus pilot
median Focus judgment in the pilot
changes to your agent's behavior in shadow mode
Your host still runs the tools, owns permissions, and shows approvals. Jev, a small fast model, answers one bounded question at a time about evidence the host hands it.
Loads only the tools and skills that apply, so prompt cost tracks the task, not your catalog.
Decides whether a command or write runs on its own, waits for you, or never runs.
Withholds tool output that carries planted instructions. The main model never reads it.
Keeps the sections the task needs. Everything else stays recoverable on request.
Checks the result actually satisfies what was asked.
Notices the same failure repeating and asks for a new approach.
Routes routine steps to a cheaper model when that is enough.
Counts, exit codes, durations, budgets, and path containment are computed in code, never by a model. A write outside the project root always needs your approval, and system paths are refused outright.
Each task asks a question about real Vitest, tsc, git log or ripgrep output. RTK is an open-source CLI that shrinks command output with fixed filters (the real v0.49.0 binary). Characters are totals across all 12 tasks, measured on the shipped Focus. Accuracy comes from the Focus pilot on the same tasks, 2 runs each. Both brainstem misses were exhaustive counts across every file; counting now goes to code.
Every number came from real captures and real API calls. Pilot results are labeled as such.
Shadow mode runs every judgment and logs what it would have done, without applying any verdicts. Output is still capped at 8,000 characters and tagged with a completeness note. Works with Pi agents today; other hosts can call the reflexes directly.
bun add @brainstem/reflexes @brainstem/pi-adapter
export TYPESAFE_API_KEY=…import { createReflexes, jevJudge } from '@brainstem/reflexes'
import { attachReflexes } from '@brainstem/pi-adapter'
const reflexes = createReflexes({ judge: jevJudge() }) // reads TYPESAFE_API_KEY
// agent: your existing Pi agent
const plugin = attachReflexes(agent, reflexes, {
cwd: process.cwd(),
modes: { gate: 'shadow', sanitize: 'shadow', verify: 'shadow' },
})
await plugin.prompt('Fix the failing login test')By default Gate, Sanitize and Verify are active and the other reflexes are off. The agent.ts example puts all three in shadow so you can watch first.