Bantam: tiny, powerful, DIY AI agent
About
Bantam is a minimalist, dependency-free AI agent specification with reference implementations in Python (bantam.py, ~280 SLOC), Go (main.go + term_*.go, module code.luxferre.top/luxferre/bantam), and Perl 5 (bantam.pl, ~360 SLOC). It provides an agentic loop capable of autonomous tool execution, shell interaction, real-time response streaming, Fibonacci backoff network resilience, and subagent delegation using any OpenAI-compatible completions API.
The entire philosophy of Bantam is built upon two principles:
- The structure must be as simple as possible for anyone to be able to reimplement the agent from a plain algorithm description.
- The agent only needs to provide two tools: a tool to call shell commands and a tool to call itself. In theory, this should be sufficient to give LLMs the ability to handle tasks of any complexity.
Because of the second principle, Bantam itself was named after Victorinox Bantam Alox, a small and lightweight Swiss army knife with only two tools.
Usage
Prerequisites
- Python 3.7+, Go 1.21+, or Perl 5.14+ (standard library / core modules only)
- An OpenAI-compatible API endpoint (or OpenAI API key)
Installation (Go)
go install code.luxferre.top/luxferre/bantam@latest
This installs the bantam binary into $(go env GOPATH)/bin (make sure it is on your PATH). To build from a local checkout instead:
go build ./... # produces ./bantam
# or run without building:
go run . prompt.txt
The Go port is a single main.go plus four platform files (term_linux.go, term_bsd.go, term_windows.go, term_other.go) for the built-in raw-terminal line editor — zero external dependencies, same as the Python and Perl versions.
Running Bantam
All implementations read the same model.cfg and system.txt from the current working directory.
-
Configure
model.cfgwith your API settings:endpoint=https://opencode.ai/zen/v1 model=deepseek-v4-flash-free temperature=0.7 api_key=your_api_key_here stream=true -
Interactive mode:
bantam # Go (or: go run .) python3 bantam.py # Python perl bantam.pl # Perl 5In interactive mode, prompts can span multiple lines: press Ctrl+J to insert a real line break (the cursor moves to the next line), then Enter to submit the whole multi-line prompt. The Go port ships its own raw-mode line editor (arrow keys move the cursor, Up/Down browse history, Backspace edits, Ctrl+C/Ctrl+D exit), so this works everywhere without dependencies; the Python and Perl ports use
readlinewhen available and fall back to single-line prompts otherwise.Sessions are saved under
~/.bantam/sessions/and can be managed with these commands:/save— save the entire conversation to a new session file (auto-id like20260808-190038) and generate its summary/list— list saved sessions (newest first) with their ids, timestamps, message counts and summaries/load <id>— load a saved session (exact id or unique prefix) and continue from there/compact— summarize the conversation with the LLM and compact the context down to just the system message plus the summary/cfg <param> [val]— inspect or update a configuration parameter inmodel.cfglive/help— show all supported commands/clear— reset the conversation to just the system prompt/quit— exit
The current conversation is also auto-saved to
~/.bantam/sessions/autosave.jsonafter every turn, on/clear,/load,/compact, and on exit — so you can always/load autosaveto resume where you left off. -
File input mode:
python3 bantam.py prompt.txt # Python bantam prompt.txt # Go perl bantam.pl prompt.txt # Perl 5
Rules of Bantam (The Algorithm)
Using these rules, everyone can build their own copy of Bantam from scratch in little time in any language that supports file access, shell and HTTP(S) calls.
High-Level Overview
- Initialization: Read
system.txtandmodel.cfg. Prepare an array of messages starting with the system prompt{"role": "system", "content": system_prompt}. - Input Processing: Take user prompt (via command-line file parameter or interactive stdin), append
{"role": "user", "content": prompt}, and invokeAL(cfg, messages). - Agentic Loop (
AL):- Send
messagesand tool definitions to the OpenAI-compatible/chat/completionsAPI endpoint with customUser-Agentheaders. - On network or HTTP failure, retry using Fibonacci backoff delays (
1s, 1s, 2s, 3s, 5s, 8s, 13s, 21s, 34s). - If
stream=true, parse SSE data chunks (data: {...}) in real-time to stream reasoning content (reasoning_content) and response text directly to stdout, bracketing the reasoning block with--- reasoning start ---/--- reasoning end ---markers. - Reconstruct the assistant message. If
tool_callsexist, trace the call ([tool call: name(args)]), validate JSON arguments, execute the requested tool (shell_execorrun_subagent), trace the result ([tool result: name]), append the tool response{"role": "tool", "tool_call_id": id, "content": result}, and repeat the loop. - If no tool calls remain or
max_al_iterationsis reached, return the updated messages list.
- Send
Main program
- Read system prompt from
system.txt(default if missing). - Read model parameters from
model.cfg(key=valueformat). - Prepare a new message list with the system prompt (
role: "system"). - Read the first command-line parameter. If non-empty, read user prompt from the specified file, append to
messages(role: "user"), runAL(cfg, messages), and exit. - Read user prompt from standard input (with
readlineline editing and history in~/.bantam_history; Ctrl+J inserts a real newline into the line being edited). If equal to/quitor EOF, exit. If equal to/clear, resetmessagesto step 3 and return to step 5. If equal to/save, write the wholemessagesarray to~/.bantam/sessions/<id>.json(with an auto-generated summary) and return to step 5. If equal to/list, print saved sessions and their summaries and return to step 5. If starting with/load, replacemessageswith the saved session's messages (by exact id or unique prefix) and return to step 5. If equal to/compact, ask the LLM to summarize the conversation, replacemessageswith[system, summary-user-message], and return to step 5. If starting with/cfg, display the current value (/cfg <param>) or updatemodel.cfglive (/cfg <param> <val>) and return to step 5. If equal to/help, print the command list and return to step 5. After every user turn and on exit, auto-savemessagesto~/.bantam/sessions/autosave.json. - Append user prompt to
messages(role: "user"), runAL(cfg, messages), and go to step 5.
Agentic loop (AL(cfg, messages)) function
- Call OpenAI-compatible Completions API (
POST {endpoint}/chat/completions) forwarding relevant parameters fromcfg(model,temperature,stream,reasoning_effort, etc., excluding internal agent configs likeendpoint,api_key,timeout,shell_timeout,max_al_iterations,color) and optionalapi_keybearer header.- Set custom
User-Agentheader (Mozilla/5.0 (compatible; Bantam/1.0)) to avoid gateway 403 blocks. - Retry network/HTTP errors with Fibonacci backoff delays (
1s, 1s, 2s, 3s, 5s, 8s, 13s, 21s, 34s). - If
stream=true, parse SSE stream (data: {...}) for real-time reasoning and text output, bracketing reasoning with--- reasoning start ---/--- reasoning end ---markers.
- Set custom
- Append the assistant's response message object to
messages. If non-streaming and response has reasoning tokens (reasoning_contentorreasoning), output them wrapped in--- reasoning start ---/--- reasoning end ---markers. - If there are pending
tool_callsin the assistant response:- For each tool call, output a trace log (
[tool call: name(args)]). - Validate JSON arguments. If invalid, format a tool-error response so the LLM can self-correct.
- Execute tool action (
shell_execorrun_subagent). - Output a trace log of the result (
[tool result: name]). - Append tool result message (
role: "tool",tool_call_id,content: result string) tomessages. - Loop back to step 1.
- For each tool call, output a trace log (
- If no pending tool calls (or if
max_al_iterationsis reached), stop and returnmessages.
If the API rejects the request with an Invalid assistant message: content or tool_calls must be set error (usually caused by a previously cut-off stream that left an empty assistant message in the session), all implementations strip the last assistant-role message from the session and retry the call.
Model configuration parameters
(shared by all implementations; model.cfg is plain key=value with # comments)
endpoint(base OpenAI-compatible API URL, defaulthttps://opencode.ai/zen/v1)model(model name, e.g.gpt-4o)temperature(model temperature, default 0.7)api_key(API key / Bearer token, optional; fall back toOPENAI_API_KEYenv var)stream(stream response tokens in real-time, defaulttrue)color(ANSI coloring:auto(TTY-detected, default),always, ornever; also disabled byNO_COLOR/BANTAM_NO_COLORenv vars)timeout(HTTP timeout in seconds for LLM API calls, default 300; in the Go port it bounds connection setup and time-to-first-byte, so long streaming responses are not cut off mid-stream, matching the Python and Perl ports' per-operation socket timeouts)shell_timeout(timeout in seconds forshell_execcommands, default 120)max_al_iterations(max tool-call loop iterations perAL()invocation, default 1000)
All implementations keep the interactive prompt safe against the classic "long line overwrites the prompt" readline bug: the Python and Perl ports wrap the ANSI escapes in \001/\002 (RL_PROMPT_START_IGNORE/RL_PROMPT_END_IGNORE) markers (disabling Term::ReadLine ornaments in Perl), and the Go port's built-in editor tracks the cursor with its own column math (terminal auto-wrap aware) and redraws from the first line of the buffer, so wrapped input stays clean at any terminal width.
When enabled, the interactive console uses a subtle ANSI palette: the pending-request status ...requesting... is darkened bold, reasoning markers are cyan, reasoning text is dim, [tool call: ...] traces are yellow, [tool result: ...] headers are green, tool result bodies are dim (red for tool errors/unknown tools), and errors/network retries are red. Tool result payloads fed back to the LLM are never colored.
Tool call definitions
run_subagent tool
- Parameters:
prompt(string) - Return value: string
- Action: run
AL(cfg, [{"role": "system", "content": system_prompt + "\n\nImportant: this is a child agent"}, {"role": "user", "content": prompt}])subject to recursion depth limit (MAX_DEPTH = 5) and return the text content of the lastassistant-role message.
shell_exec tool
- Parameters:
command(string) - Return value: string
- Action: run shell command specified in
commandsubject toshell_timeout(default 120s) and returnoutput + '\n\nexit: ' + exit_codestring.
MicroBantam
MicroBantam (mb) is a compressed Perl 5 reference implementation of the same agent in under 100 SLOC, written to stay readable while keeping the full agentic core. It reads the same model.cfg and system.txt from the current working directory.
Features
- Full agentic loop: LLM calls,
shell_exec/run_subagenttool execution with JSON argument validation (invalid args are fed back so the model can self-correct), and the 5-level subagent recursion depth limit - A
...requesting...in-flight indicator: in-place on a TTY (\roverwrite, erased on completion), a plain line when output is piped - Session management:
/save,/list,/load <id>(exact id only, no prefix matching), and auto-save to~/.bantam/sessions/autosave.jsonafter every turn and on exit; session ids get-1,-2, ... suffixes on same-second collisions - Interactive mode (
/quit,/clear,/save,/list,/load <id>,/help) and file input mode - Same built-in default system prompt and
OPENAI_API_KEYfallback as the full implementation - Automatic recovery from
Invalid assistant message: content or tool_calls must be setAPI errors: strips the last assistant message and retries
What it drops
- Streaming (requests are non-streaming;
streamis ignored) - ANSI coloring/styling (
coloris ignored) Term::ReadLineline editing, Ctrl+J multi-line prompts and readline history (plain single-line prompts)- Fibonacci backoff network retries (a failed request aborts with an
API errormessage) /compactcontext summarization
Jim Tcl port
mb.tcl is a highly experimental Jim Tcl port of the same agent (under 100 SLOC) with the same feature set as the Perl mb. It differs in three ways: it ships its own minimal HTTP client (raw sockets with chunked-transfer decoding) instead of HTTP::Tiny; it retries failed requests up to 3 times with a 2-second backoff instead of aborting on the first failure; and it relies on the external timeout command for shell_exec timeouts instead of SIGALRM. Run it the same way:
./mb.tcl # interactive (or: jimsh mb.tcl)
./mb.tcl prompt.txt # file input mode
Running
./mb # interactive (or: perl mb)
./mb prompt.txt # file input mode
Repository layout
bantam.py— Python reference implementation (stdlib only)main.go,term_linux.go,term_bsd.go,term_windows.go,term_other.go— Go implementation (stdlib only, modulecode.luxferre.top/luxferre/bantam)bantam.pl— Perl 5 implementation (core modules only)mb— MicroBantam, compressed Perl 5 implementation (core modules only, under 100 SLOC)mb.tcl— MicroBantam, Jim Tcl port (requiresjimshwith thejsonandsslextensions; under 100 SLOC)model.cfg,system.txt— shared configuration and system promptREADME.md— this document
FAQ
Does Bantam support AGENTS.md etc?
The default system prompt instructs the agent to respect AGENTS.md/GEMINI.md/CLAUDE.md files.
Is there any common config place for Bantam?
No, loading model.cfg and system.txt is deliberately only supported from the current working directory. This allows natural separation of configs and system prompts per project. In case there's no system.txt inside the project, the concise and sensible default system prompt will be loaded. In case there's no model.cfg inside the project, Bantam will use the free Big Pickle model from OpenCode Zen with the temperature 0.7. Big Pickle has been chosen as the default because it has no set expiration date, unlike other OpenCode's keyless tiers.
Why no MCP support?
If you need MCP server tools, there's nothing a simple shell wrapper cannot solve in this case.
Why no sandboxing?
Same philosophy as the Pi agent: there's nothing a simple chroot environment cannot solve in case sandboxing is really necessary.
Any advanced authentication schemes or header injection?
You can pair Bantam with the Dynagate LLM gateway to achieve all that.
How to run on mobiles?
On Android, any current Bantam/MicroBantam implementation is easily runnable within the Termux environment. Go implementation is preferred for performance reasons.
On iOS/iPadOS, the easiest way to use Bantam is to run the Perl version (or MicroBantam) inside iSH. Some terminal features may not be available (run with rlwrap to bring them back), but the agent itself is fully functional.
Credits
Created by Luxferre in 2026, released into the public domain with no warranties.