[00] agent runtime
Agents that ship work, not demos.
Runlog is the runtime for production AI agents. Plan, call tools, verify the result, and replay every step when something looks off.
$ npm i @runlog/sdk- 01split "flaky checkout test" into 3 tasks0.42s
- 02repo.search(q: "checkout timeout")1.18s
- 03ci.logs(run: 4821, tail: 200)0.87s
- 04root cause: cache warmup race2.03s
- 05repo.create_pr(branch: "fix/cache-race")1.36s
- 06tests 48/48 passed, score 0.970.64s
tokens14,212
cost$0.031
steps6/6
[01] running in production at
- northwind
- Halcyon
- kestrel/ai
- AMPERE
- fieldnote
- Quarry
- Ashton
[02] the loop
Plan. Act. Verify. On every run.
Each step is explicit, budgeted and checked. A bad step stops the run instead of reaching your users.
- 01PLAN
PLAN
Plans before it acts
Typed plan, step and cost budget.
plan 3 steps · $0.10
- 02ACT
ACT
Calls tools you trust
Schema-checked, idempotent, approvals.
tool repo.create_pr 412ms
- 03VERIFYSHIP ✓
VERIFY
Checks its own work
Evals gate every result.
eval 0.97 >= 0.90 pass
- SHIP ✓
[03] sdk
Twelve lines to a production agent.
Declare tools, a budget and a verify step. Runlog handles retries, streaming, approvals and the audit trail.
- +Typed tools in TypeScript or Python
- +Streaming steps over SSE or webhooks
- +Any model, swap it in one line
- +Local dev server with replays
import { agent, tool } from "@runlog/sdk";
// typed tools: inputs are validated before every call
const createPr = tool("repo.create_pr", {
input: { branch: "string", title: "string" },
approval: "required",
});
export default agent("triage-bot", {
model: "your-model",
tools: [createPr],
budget: { usd: 0.10, steps: 12 },
verify: async (run) => run.tests.passed,
});from runlog import agent, tool
# typed tools: inputs are validated before every call
create_pr = tool(
"repo.create_pr",
input={"branch": str, "title": str},
approval="required",
)
triage = agent(
"triage-bot",
model="your-model",
tools=[create_pr],
budget={"usd": 0.10, "steps": 12},
verify=lambda run: run.tests.passed,
)# start a run and stream every step back
curl https://api.example.com/v1/runs \
-H "Authorization: Bearer $RUNLOG_KEY" \
-d '{
"agent": "triage-bot",
"input": "fix the flaky checkout test",
"stream": true
}'$ runlog devrun_8f3a21 verified in 6.5s · replay at localhost:4000
[04] observability
Every run is a ledger.
Know what every agent did, why, and what it cost.
- >>Replay any stepRe-run from any point with new inputs.
- +-Diff two runsSee exactly where behavior changed.
- ->Export anywhereStream runs to your warehouse.
| run | agent | status | steps |
|---|---|---|---|
| run_8f3a21 | triage-bot | passed | 6.5s |
| run_8f3a1c | refund-desk | passed | 3.4s |
| run_8f39e7 | lead-router | retried | 6.6s |
| run_8f39b0 | triage-bot | passed | 4.9s |
| run_8f3982 | docs-answer | passed | 1.3s |
[05] in numbers
- 41M
- tool calls a month
- 99.95%
- uptime, backed by SLA
- 38ms
- median overhead
- 0
- unapproved writes
[06] tools
120+ tools ready. Or write your own in ten lines.
- PRrepo.create_propen pull requests
- #chat.postmessage a channel
- DBsql.queryread-only by default
- $payments.refundapproval required
- ?docs.searchhybrid retrieval
- @mail.sendtemplated email
- //browser.openheadless sessions
- >_python.execsandboxed runtime
- {}http.requestany REST API
- <>mcp.connectany MCP server
- 31calendar.bookfind a free slot
- +tool.custombring your own
[07] guardrails
Built for the audit, not just the demo.
Security reviews go faster when the answers are already in the product.
! approval required
refund-desk wants to call
payments.refund(order: 1182, amount: $240)
Human approvals
approval: requiredRuns pause until someone approves in chat or the dashboard.
Scoped credentials
scope: repo:writeNarrowest keys per agent. Secrets never enter the context.
Full audit trail
retention: 400dEvery prompt, call and result, searchable.
Self-host or pin a region
region: eu-centralYour cloud, or data kept in the EU or US.
[08] from the field
"We moved four agents from a notebook to production in two weeks. The replay view alone saved us a hire."
- 4
- agents in production
- 2 wks
- notebook to prod
- 0
- incidents since launch
[10] exit 0
Start your first run.
Free for 10,000 steps a month. No card, no sales call.
$ runlog deploy triage-bot live in 4.2s