Files
k3r053n3/README.md
T
2026-09-02 16:19:51 +03:00

308 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# k3r053n3: Kimi K3 OpenAI-compatible proxy gateway
A standalone, high-performance, zero-dependency Go proxy gateway that exposes a standard OpenAI-compatible API (`/v1/chat/completions` and `/v1/models`) for the **Kimi K3** model hosted on the [`cw-105/kimi-k3-gguf-demo`](https://cw-105-kimi-k3-gguf-demo.hf.space) Gradio space.
## Features
- **Zero external dependencies**: pure Go standard library (`net/http`, `encoding/json`, `bufio`, `bytes`, `crypto/rand`, `flag`, `strings`, `time`).
- **SOCKS5 proxy support**: zero-dependency built-in RFC 1928 / RFC 1929 SOCKS5 client supporting domain name resolution (`socks5h://`), IPv4/IPv6, and username/password authentication (via `-socks`, `-proxy`, or `ALL_PROXY`/`all_proxy`/`SOCKS5_PROXY`/`socks5_proxy` environment variables).
- **Reasoning effort support**: defaults to `"max"` reasoning effort out of the box, with support for `"max"`, `"high"`, `"low"`, and `"default"` (via `reasoning_effort` request field or CLI flag).
- **Real-time token streaming**: Server-Sent Events (SSE) streaming engine (`stream: true`) with separate token-by-token emission for `delta.reasoning_content` and `delta.content`.
- **Reasoning extraction**: clean separation of `<think>...</think>` internal thoughts into `reasoning_content` (streaming chunks and non-streaming messages) without leaking raw tags into `content`.
- **OpenAI-compatible tool calling**:
- automatic tool definition formatting into system instructions.
- multi-turn tool execution response formatting (`role: "tool"` / `role: "function"`).
- real-time stream interceptor (`StreamToolInterceptor`) that catches `<tool_call>` blocks on the fly and emits standard OpenAI `delta.tool_calls` chunks with `finish_reason: "tool_calls"`.
- non-streaming tool call parsing with structured `tool_calls` and `finish_reason: "tool_calls"`.
- **Backend & model routing**: route between multiple upstream backends (`direct:together`, `direct:fireworks`, `hf:together`, `hf:fireworks-ai`, `hf:featherless-ai`, `hf:baseten`) dynamically or via model suffix (`kimi-k3:together`, `kimi-k3:fireworks`, etc.).
- **Reliability & resilience**: automatic Fibonacci exponential backoff retry mechanism (`DoWithFibonacciRetry`) on upstream network connections.
- **Full CORS support**: ready for direct browser integration, web frontends, and OpenAI-compatible client libraries.
## Quick start
### Installation
Install directly with `go install`:
```bash
go install code.luxferre.top/luxferre/k3r053n3@latest
```
### Build from source
```bash
make k3r053n3
```
Or build manually with Go:
```bash
go build -trimpath -ldflags="-s -w" -o bin/k3r053n3 .
```
### Run
```bash
./bin/k3r053n3
```
By default, the server starts on port `8080` pointing to `https://cw-105-kimi-k3-gguf-demo.hf.space` with default reasoning effort `"max"` and default backend `"direct:together"`.
## CLI options
| Flag | Default | Description |
||||
| `-port` | `8080` | Port to listen on |
| `-endpoint` | `https://cw-105-kimi-k3-gguf-demo.hf.space` | Root URL of the Kimi K3 Gradio space |
| `-model` | `kimi-k3` | Exposed default model name |
| `-backend` | `direct:together` | Default Space backend (`direct:together`, `direct:fireworks`, `hf:together`, `hf:fireworks-ai`, `hf:featherless-ai`, `hf:baseten`) |
| `-reasoning` | `max` | Default reasoning effort (`max`, `high`, `low`, `default`) |
| `-max-tokens` | `8192` | Default max completion tokens (256 - 8192) |
| `-temperature` | `0.7` | Default sampling temperature (0.0 - 1.5) |
| `-socks`, `-proxy` | `""` | SOCKS5 proxy URL (`socks5://127.0.0.1:1080` or `socks5://user:pass@host:port`, also checks `ALL_PROXY`/`all_proxy`/`SOCKS5_PROXY`/`socks5_proxy` env vars) |
| `-user-agent`, `-ua` | Firefox 153 on Linux | Custom `User-Agent` header for upstream requests |
## Endpoints
- `GET /` - gateway health check and route overview
- `GET /v1/models` (or `GET /models`) - list of available models and backend mappings
- `POST /v1/chat/completions` (or `POST /chat/completions`) - OpenAI-compatible chat completions
## Usage examples
### 1. Models list
```bash
curl -s http://localhost:8080/v1/models
```
### 2. Standard chat completion (non-streaming)
```bash
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "What is 25 * 4? Show brief work."}
],
"stream": false
}'
```
Response includes separated `reasoning_content` and `content`:
```json
{
"id": "chatcmpl-88054d2e-2084-4bcd-b9fa-8e99e9267523",
"object": "chat.completion",
"created": 1788348384,
"model": "kimi-k3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "**25 × 4 = 100**\n\nQuick way: 25 × 4 = 25 × 2 × 2 = 50 × 2 = **100**\n\n(Think of it as 4 quarters = 1 dollar.)",
"reasoning_content": "The user is asking a simple arithmetic question: 25 * 4..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
```
### 3. Real-time streaming with reasoning deltas
```bash
curl -s -N http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "Tell me a 1-sentence joke about computers."}
],
"stream": true
}'
```
Output delivers real-time `delta.reasoning_content` chunks during thinking, followed by `delta.content` chunks for the answer, ending with `[DONE]`:
```
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1788348397,"model":"kimi-k3","choices":[{"index":0,"delta":{"reasoning_content":"Thinking..."}}]}
...
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1788348397,"model":"kimi-k3","choices":[{"index":0,"delta":{"content":"There are only 10 types of people in the world: those who understand binary and those who don't."}}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1788348397,"model":"kimi-k3","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
```
### 4. Reasoning effort control
Control reasoning effort via the `reasoning_effort` field (`"max"`, `"high"`, `"low"`, `"default"`):
```bash
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"reasoning_effort": "low",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
```
### 5. Tool / function calling
```bash
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "What is the weather in Seattle right now?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a given city",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}
],
"stream": false
}'
```
Response emits standard OpenAI `tool_calls` with `finish_reason: "tool_calls"`:
```json
{
"id": "chatcmpl-cbe006c5-1b5d-42d3-ab89-254c7f012061",
"object": "chat.completion",
"created": 1788348416,
"model": "kimi-k3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"reasoning_content": "The user is asking about the weather in Seattle...",
"tool_calls": [
{
"index": 0,
"id": "call_de6cdf1a_0",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\":\"Seattle\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
]
}
```
### 6. Submitting tool results in follow-up turns
```bash
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "What is the weather in Seattle right now?"},
{
"role": "assistant",
"tool_calls": [
{
"id": "call_de6cdf1a_0",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"location\":\"Seattle\"}"}
}
]
},
{
"role": "tool",
"tool_call_id": "call_de6cdf1a_0",
"name": "get_weather",
"content": "{\"temperature\": \"16C\", \"conditions\": \"Partly cloudy with gentle breeze\"}"
}
]
}'
```
## Python OpenAI client integration
```python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8080/v1",
api_key="not-needed",
)
# Streaming with reasoning
response = client.chat.completions.create(
model="kimi-k3",
messages=[
{"role": "user", "content": "Explain quantum superposition in 2 sentences."}
],
stream=True,
extra_body={"reasoning_effort": "max"},
)
for chunk in response:
delta = chunk.choices[0].delta
if hasattr(delta, "reasoning_content") and delta.reasoning_content:
print(delta.reasoning_content, end="", flush=True)
if delta.content:
print(delta.content, end="", flush=True)
print()
```
## SOCKS5 proxy usage
Run the gateway through a SOCKS5 proxy (e.g. Tor or local tunnel):
```bash
# Using CLI flag
./bin/k3r053n3 -socks socks5://127.0.0.1:9050
# With authentication
./bin/k3r053n3 -socks socks5://user:pass@127.0.0.1:1080
# Using environment variable
export ALL_PROXY=socks5://127.0.0.1:1080
./bin/k3r053n3
```
## Testing
Run unit and integration tests:
```bash
make test
```
## Credits
Created by Luxferre in 2026, released into the public domain with no warranties.