172 lines
11 KiB
Markdown
172 lines
11 KiB
Markdown
# Bantam: tiny, powerful, DIY AI agent
|
|
|
|
## About
|
|
|
|
Bantam is a minimalist, dependency-free AI agent specification with two reference implementations: **Python** (`bantam.py`, ~300 SLOC) and **Go** (`main.go` + `term_*.go`, module `code.luxferre.top/luxferre/bantam`). It provides an agentic loop capable of autonomous tool execution, shell interaction, real-time response streaming, Fibonacci backoff network resilience, and subagent delegation using any OpenAI-compatible completions API.
|
|
|
|
The entire philosophy of Bantam is built upon two principles:
|
|
|
|
1. The structure must be as simple as possible for anyone to be able to reimplement the agent from a plain algorithm description.
|
|
2. The agent only needs to provide two tools: a tool to call shell commands and a tool to call itself. In theory, this should be sufficient to give LLMs the ability to handle tasks of any complexity.
|
|
|
|
Because of the second principle, Bantam itself was named after Victorinox Bantam Alox, a small and lightweight Swiss army knife with only two tools.
|
|
|
|
## Usage
|
|
|
|
### Prerequisites
|
|
- **Go 1.21+** (Go implementation, no external dependencies) or **Python 3.7+** (reference implementation)
|
|
- An OpenAI-compatible API endpoint (or OpenAI API key)
|
|
|
|
### Installation (Go)
|
|
|
|
```bash
|
|
go install code.luxferre.top/luxferre/bantam@latest
|
|
```
|
|
|
|
This installs the `bantam` binary into `$(go env GOPATH)/bin` (make sure it is on your `PATH`). To build from a local checkout instead:
|
|
|
|
```bash
|
|
go build ./... # produces ./bantam
|
|
# or run without building:
|
|
go run . prompt.txt
|
|
```
|
|
|
|
The Go port is a single `main.go` plus two small platform files (`term_linux.go`, `term_darwin.go`, `term_windows.go`, `term_other.go`) for the built-in raw-terminal line editor — zero external dependencies, same as the Python version.
|
|
|
|
### Running Bantam
|
|
|
|
Both implementations read the same `model.cfg` and `system.txt` from the current working directory.
|
|
|
|
1. Configure `model.cfg` with your API settings:
|
|
```ini
|
|
endpoint=https://api.openai.com/v1
|
|
model=gpt-4o
|
|
temperature=0.7
|
|
api_key=your_api_key_here
|
|
stream=true
|
|
```
|
|
|
|
2. Interactive mode:
|
|
```bash
|
|
bantam # Go (or: go run .)
|
|
python3 bantam.py # Python
|
|
```
|
|
|
|
In interactive mode, prompts can span multiple lines: press **Ctrl+J** to insert a real line break (the cursor moves to the next line), then **Enter** to submit the whole multi-line prompt. The Go port ships its own raw-mode line editor (arrow keys move the cursor, Up/Down browse history, Backspace edits, Ctrl+C/Ctrl+D exit), so this works everywhere without dependencies; the Python port uses `readline` when available and falls back to single-line prompts otherwise.
|
|
|
|
Sessions are saved under `~/.bantam/sessions/` and can be managed with these commands:
|
|
- `/save` — save the entire conversation to a new session file (auto-id like `20260808-190038`) and generate its summary
|
|
- `/list` — list saved sessions (newest first) with their ids, timestamps, message counts and summaries
|
|
- `/load <id>` — load a saved session (exact id or unique prefix) and continue from there
|
|
- `/compact` — summarize the conversation with the LLM and compact the context down to just the system message plus the summary
|
|
- `/help` — show all supported commands
|
|
- `/clear` — reset the conversation to just the system prompt
|
|
- `/quit` — exit
|
|
|
|
The current conversation is also **auto-saved** to `~/.bantam/sessions/autosave.json` after every turn, on `/clear`, `/load`, `/compact`, and on exit — so you can always `/load autosave` to resume where you left off.
|
|
|
|
3. File input mode:
|
|
```bash
|
|
python3 bantam.py prompt.txt # Python
|
|
bantam prompt.txt # Go
|
|
```
|
|
|
|
## Rules of Bantam (The Algorithm)
|
|
|
|
Using these rules, everyone can build their own copy of Bantam from scratch in little time in any language that supports file access, shell and HTTP(S) calls.
|
|
|
|
### High-Level Overview
|
|
|
|
1. **Initialization**: Read `system.txt` and `model.cfg`. Prepare an array of messages starting with the system prompt `{"role": "system", "content": system_prompt}`.
|
|
2. **Input Processing**: Take user prompt (via command-line file parameter or interactive stdin), append `{"role": "user", "content": prompt}`, and invoke `AL(cfg, messages)`.
|
|
3. **Agentic Loop (`AL`)**:
|
|
- Send `messages` and tool definitions to the OpenAI-compatible `/chat/completions` API endpoint with custom `User-Agent` headers.
|
|
- On network or HTTP failure, retry using Fibonacci backoff delays (`1s, 1s, 2s, 3s, 5s`).
|
|
- If `stream=true`, parse SSE data chunks (`data: {...}`) in real-time to stream reasoning content (`reasoning_content`) and response text directly to stdout, bracketing the reasoning block with `--- reasoning start ---` / `--- reasoning end ---` markers.
|
|
- Reconstruct the assistant message. If `tool_calls` exist, trace the call (`[tool call: name(args)]`), validate JSON arguments, execute the requested tool (`shell_exec` or `run_subagent`), trace the result (`[tool result: name]`), append the tool response `{"role": "tool", "tool_call_id": id, "content": result}`, and repeat the loop.
|
|
- If no tool calls remain or `max_al_iterations` is reached, return the updated messages list.
|
|
|
|
---
|
|
|
|
### Main program
|
|
|
|
1. Read system prompt from `system.txt` (default if missing).
|
|
2. Read model parameters from `model.cfg` (`key=value` format).
|
|
3. Prepare a new message list with the system prompt (`role: "system"`).
|
|
4. Read the first command-line parameter. If non-empty, read user prompt from the specified file, append to `messages` (`role: "user"`), run `AL(cfg, messages)`, and exit.
|
|
5. Read user prompt from standard input (with `readline` line editing and history in `~/.bantam_history`; **Ctrl+J** inserts a real newline into the line being edited). If equal to `/quit` or EOF, exit. If equal to `/clear`, reset `messages` to step 3 and return to step 5. If equal to `/save`, write the whole `messages` array to `~/.bantam/sessions/<id>.json` (with an auto-generated summary) and return to step 5. If equal to `/list`, print saved sessions and their summaries and return to step 5. If starting with `/load`, replace `messages` with the saved session's messages (by exact id or unique prefix) and return to step 5. If equal to `/compact`, ask the LLM to summarize the conversation, replace `messages` with `[system, summary-user-message]`, and return to step 5. If equal to `/help`, print the command list and return to step 5. After every user turn and on exit, auto-save `messages` to `~/.bantam/sessions/autosave.json`.
|
|
6. Append user prompt to `messages` (`role: "user"`), run `AL(cfg, messages)`, and go to step 5.
|
|
|
|
### Agentic loop (`AL(cfg, messages)`) function
|
|
|
|
1. Call OpenAI-compatible Completions API (`POST {endpoint}/chat/completions`) using parameters from `cfg` (`model`, `temperature`, optional `api_key` bearer header).
|
|
- Set custom `User-Agent` header (`Mozilla/5.0 (compatible; Bantam/1.0)`) to avoid gateway 403 blocks.
|
|
- Retry network/HTTP errors with Fibonacci backoff delays (`1s, 1s, 2s, 3s, 5s`).
|
|
- If `stream=true`, parse SSE stream (`data: {...}`) for real-time reasoning and text output, bracketing reasoning with `--- reasoning start ---` / `--- reasoning end ---` markers.
|
|
2. Append the assistant's response message object to `messages`. If non-streaming and response has reasoning tokens (`reasoning_content` or `reasoning`), output them wrapped in `--- reasoning start ---` / `--- reasoning end ---` markers.
|
|
3. If there are pending `tool_calls` in the assistant response:
|
|
- For each tool call, output a trace log (`[tool call: name(args)]`).
|
|
- Validate JSON arguments. If invalid, format a tool-error response so the LLM can self-correct.
|
|
- Execute tool action (`shell_exec` or `run_subagent`).
|
|
- Output a trace log of the result (`[tool result: name]`).
|
|
- Append tool result message (`role: "tool"`, `tool_call_id`, `content`: result string) to `messages`.
|
|
- Loop back to step 1.
|
|
4. If no pending tool calls (or if `max_al_iterations` is reached), stop and return `messages`.
|
|
|
|
### Model configuration parameters
|
|
|
|
(shared by both implementations; `model.cfg` is plain `key=value` with `#` comments)
|
|
|
|
- `endpoint` (base OpenAI-compatible API URL, default `https://api.openai.com/v1`)
|
|
- `model` (model name, e.g. `gpt-4o`)
|
|
- `temperature` (model temperature, default 0.7)
|
|
- `api_key` (API key / Bearer token, optional; fall back to `OPENAI_API_KEY` env var)
|
|
- `stream` (stream response tokens in real-time, default `true`)
|
|
- `color` (ANSI coloring: `auto` (TTY-detected, default), `always`, or `never`; also disabled by `NO_COLOR`/`BANTAM_NO_COLOR` env vars)
|
|
- `timeout` (HTTP timeout in seconds for LLM API calls, default 60)
|
|
- `shell_timeout` (timeout in seconds for `shell_exec` commands, default 120)
|
|
- `max_al_iterations` (max tool-call loop iterations per `AL()` invocation, default 1000)
|
|
|
|
Both implementations keep the interactive prompt safe against the classic "long line overwrites the prompt" readline bug: the Python port wraps the ANSI escapes in `\001`/`\002` (`RL_PROMPT_START_IGNORE`/`RL_PROMPT_END_IGNORE`) markers, and the Go port's built-in editor tracks the cursor with its own column math (terminal auto-wrap aware) and redraws from the first line of the buffer, so wrapped input stays clean at any terminal width.
|
|
|
|
When enabled, the interactive console uses a subtle ANSI palette: the pending-request status `...requesting...` is darkened bold, reasoning markers are cyan, reasoning text is dim, `[tool call: ...]` traces are yellow, `[tool result: ...]` headers are green, tool result bodies are dim (red for tool errors/unknown tools), and errors/network retries are red. Tool result payloads fed back to the LLM are never colored.
|
|
|
|
### Tool call definitions
|
|
|
|
#### `run_subagent` tool
|
|
|
|
- Parameters: `prompt` (string)
|
|
- Return value: string
|
|
- Action: run `AL(cfg, [{"role": "system", "content": system_prompt + "\n\nImportant: this is a child agent"}, {"role": "user", "content": prompt}])` and return the text content of the last `assistant`-role message.
|
|
|
|
#### `shell_exec` tool
|
|
|
|
- Parameters: `command` (string)
|
|
- Return value: string
|
|
- Action: run shell command specified in `command` subject to `shell_timeout` (default 120s) and return `output + '\n\nexit: ' + exit_code` string.
|
|
|
|
## Repository layout
|
|
|
|
- `bantam.py` — Python reference implementation (stdlib only)
|
|
- `main.go`, `term_linux.go`, `term_darwin.go`, `term_windows.go`, `term_other.go` — Go implementation (stdlib only, module `code.luxferre.top/luxferre/bantam`)
|
|
- `model.cfg`, `system.txt` — shared configuration and system prompt
|
|
- `README.md` — this document
|
|
|
|
## FAQ
|
|
|
|
### Why no MCP support?
|
|
|
|
If you need MCP server tools, there's nothing a simple shell wrapper cannot solve in this case.
|
|
|
|
### Why no sandboxing?
|
|
|
|
Same philosophy as the Pi agent: there's nothing a simple chroot environment cannot solve in case sandboxing is really necessary.
|
|
|
|
### Any advanced authentication schemes or header injection?
|
|
|
|
You can pair Bantam with the [Dynagate](https://code.luxferre.top/luxferre/dynagate) LLM gateway to achieve all that.
|
|
|
|
## Credits
|
|
|
|
Created by Luxferre in 2026, released into the public domain with no warranties.
|