docs: use sentence case for subheadings and list items
This commit is contained in:
@@ -1,66 +1,54 @@
|
|||||||
# qflash: OpenAI Proxy Gateway for Qwen3.8-Flash-Next
|
# Qflash
|
||||||
|
|
||||||
Standalone, performant, zero-dependency Go OpenAI proxy gateway for the **Qwen3.8-Flash-Next** Hugging Face Gradio space (`https://halvo78-qwen3-8-flash-next-playground.hf.space`).
|
## About
|
||||||
|
|
||||||
Created by Luxferre in 2026, released into the public domain.
|
Qflash is a standalone, single-binary gateway that exposes the Qwen3.8-Flash-Next Gradio space (`https://halvo78-qwen3-8-flash-next-playground.hf.space`) through an OpenAI-compatible API. It translates the standard `/v1/chat/completions` and `/v1/models` endpoints into Gradio's `/gradio_api/call/chat_response` request and Server-Sent Events (SSE) stream protocol, allowing any standard OpenAI-compatible client, agent, or IDE to interface with Qwen3.8-Flash-Next without modification.
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Features
|
## Features
|
||||||
|
|
||||||
- **Zero External Dependencies**: Built entirely with Go standard library packages (`net/http`, `encoding/json`, `bufio`, etc.).
|
- OpenAI-compatible chat completions (streaming and non-streaming)
|
||||||
- **OpenAI-Compatible API**: Implements standard `/v1/chat/completions` (streaming & non-streaming) and `/v1/models`.
|
- Deep reasoning extraction with thinking trace passthrough (`<think>` tags and blockquotes mapped to `reasoning_content`)
|
||||||
- **Real-Time Token Streaming**: Streams SSE chunks with incremental token delivery directly to clients.
|
- Stateful streaming tool call interception (`StreamToolCallFilter`) with zero XML/JSON leakage into `delta.content`
|
||||||
- **Deep Reasoning Separation**:
|
- Real-time token streaming with incremental SSE delivery
|
||||||
- Automatically isolates thinking traces from both standard `<think>...</think>` tags and the playground's blockquote thinking blocks (`> 💭 **Thinking Process...**`).
|
- Support for instruct mode via standard `reasoning_effort: "none"`
|
||||||
- Emits pure thought traces to `delta.reasoning_content` (streaming) and `message.reasoning_content` (non-streaming).
|
- Zero-dependency SOCKS5 proxy client (RFC 1928 / RFC 1929) with domain resolution (`socks5h://`), IPv4, IPv6, and auth
|
||||||
- Keeps `delta.content` and `message.content` clean.
|
- Bring Your Own Key (BYOK) pass-through support via `Authorization: Bearer` or CLI flags
|
||||||
- **Stateful Streaming Tool Call Interception**:
|
- Fibonacci backoff retry on transient upstream errors
|
||||||
- Injects tool schemas into system instructions.
|
- Zero external dependencies (Go standard library only)
|
||||||
- Intercepts `<tool_call>` blocks in real time via `StreamToolCallFilter` without leaking raw XML or JSON into `delta.content`.
|
|
||||||
- Emits structured `delta.tool_calls` chunks and sets `finish_reason: "tool_calls"`.
|
|
||||||
- **Reasoning Effort Control**: Respects standard `reasoning_effort: "none"` to switch dynamically into high-speed Instruct Mode.
|
|
||||||
- **Zero-Dependency SOCKS5 Proxy Client**:
|
|
||||||
- RFC 1928 and RFC 1929 compliant client with domain resolution (`socks5h://`), IPv4, IPv6, and authentication.
|
|
||||||
- Wireable via `-socks` CLI flag or `ALL_PROXY` / `SOCKS5_PROXY` environment variables.
|
|
||||||
- **Bring Your Own Key (BYOK) Pass-through**:
|
|
||||||
- Passes client API keys or custom base URLs directly to upstream inference engines when provided.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Installation
|
## Installation
|
||||||
|
|
||||||
### Via `go install`
|
You need Go 1.22 or newer (tested on Go 1.26). Install the latest release straight from the Git repository with `go install`:
|
||||||
|
|
||||||
Install directly to `$(go env GOPATH)/bin`:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
go install code.luxferre.top/luxferre/qflash@latest
|
go install code.luxferre.top/luxferre/qflash@latest
|
||||||
```
|
```
|
||||||
|
|
||||||
*Note: If installing right after pushing to Gitea, bypass any public module proxy cache by using `GOPROXY=direct`:*
|
This fetches the module from `https://code.luxferre.top/luxferre/qflash.git` and places the `qflash` binary in `$(go env GOPATH)/bin`. Make sure that directory is on your `PATH`.
|
||||||
|
|
||||||
```bash
|
*(Note: If installing right after a new commit has been pushed, bypass any proxy cache with `GOPROXY=direct go install code.luxferre.top/luxferre/qflash@latest`).*
|
||||||
GOPROXY=direct go install code.luxferre.top/luxferre/qflash@latest
|
|
||||||
```
|
|
||||||
|
|
||||||
### From Source
|
If you prefer to build from a local checkout instead:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git clone https://code.luxferre.top/luxferre/qflash.git
|
git clone https://code.luxferre.top/luxferre/qflash.git
|
||||||
cd qflash
|
cd qflash
|
||||||
|
go install .
|
||||||
|
```
|
||||||
|
|
||||||
|
Alternatively, build directly from source using `make`:
|
||||||
|
|
||||||
|
```bash
|
||||||
make qflash
|
make qflash
|
||||||
```
|
```
|
||||||
|
|
||||||
The resulting binary will be placed at `bin/qflash`.
|
This produces the `bin/qflash` binary for your platform.
|
||||||
|
|
||||||
---
|
## Models served
|
||||||
|
|
||||||
## Architecture & Model Aliases
|
The gateway advertises the following models under `/v1/models`:
|
||||||
|
|
||||||
The gateway serves the following models under `/v1/models`:
|
| Model ID | Target model | Description |
|
||||||
|
|
||||||
| Model ID | Target Model | Description |
|
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `Qwen/Qwen3.8-Flash-Next` | `Qwen/Qwen3.8-Flash-Next` | Primary playground model (125B MoE, 6B activated) |
|
| `Qwen/Qwen3.8-Flash-Next` | `Qwen/Qwen3.8-Flash-Next` | Primary playground model (125B MoE, 6B activated) |
|
||||||
| `qwen3.8-flash-next` | `Qwen/Qwen3.8-Flash-Next` | Standard lowercase alias |
|
| `qwen3.8-flash-next` | `Qwen/Qwen3.8-Flash-Next` | Standard lowercase alias |
|
||||||
@@ -69,87 +57,62 @@ The gateway serves the following models under `/v1/models`:
|
|||||||
|
|
||||||
Any unlisted custom model name requested by the client is passed through directly.
|
Any unlisted custom model name requested by the client is passed through directly.
|
||||||
|
|
||||||
---
|
## Usage
|
||||||
|
|
||||||
## Configuration Flags & Environment Variables
|
Run the gateway:
|
||||||
|
|
||||||
| Flag | Shorthand | Environment Variable | Default | Description |
|
|
||||||
|---|---|---|---|---|
|
|
||||||
| `-port` | | `PORT` | `8080` | Port to bind the HTTP server |
|
|
||||||
| `-endpoint` | | | `https://halvo78-qwen3-8-flash-next-playground.hf.space` | Upstream Gradio space base URL |
|
|
||||||
| `-model` | | | `Qwen/Qwen3.8-Flash-Next` | Default model ID |
|
|
||||||
| `-thinking` | `-enable-thinking` | | `true` | Enable chain-of-thought reasoning by default |
|
|
||||||
| `-hf-token` | | `HF_TOKEN` | `""` | Hugging Face user access token |
|
|
||||||
| `-api-key` | | `OPENAI_API_KEY` / `QWEN_API_KEY` | `""` | Upstream inference engine API key |
|
|
||||||
| `-base-url` | | `OPENAI_BASE_URL` / `QWEN_BASE_URL` | `""` | Upstream inference engine base URL |
|
|
||||||
| `-user-agent` | `-ua` | `USER_AGENT` | Firefox 153 on Linux | Custom User-Agent header |
|
|
||||||
| `-socks` | `-proxy`, `-socks5` | `ALL_PROXY`, `SOCKS5_PROXY` | `""` | SOCKS5 proxy URL (`socks5://127.0.0.1:1080`) |
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Usage Examples
|
|
||||||
|
|
||||||
### 1. Launch Gateway
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
qflash -port 8080
|
qflash [-port 8080] [-endpoint https://halvo78-qwen3-8-flash-next-playground.hf.space] [-model Qwen/Qwen3.8-Flash-Next]
|
||||||
```
|
```
|
||||||
|
|
||||||
### 2. List Models
|
Available flags:
|
||||||
|
|
||||||
|
- `-port` — TCP port to listen on (default `8080`)
|
||||||
|
- `-endpoint` — root URL of the Gradio space (default `https://halvo78-qwen3-8-flash-next-playground.hf.space`)
|
||||||
|
- `-model` — exposed model name (default `Qwen/Qwen3.8-Flash-Next`)
|
||||||
|
- `-thinking` / `-enable-thinking` — enable chain-of-thought reasoning by default (default `true`)
|
||||||
|
- `-hf-token` — optional Hugging Face API token for authenticated spaces (`HF_TOKEN` env)
|
||||||
|
- `-api-key` — upstream inference engine API key for BYOK mode (`OPENAI_API_KEY` / `QWEN_API_KEY` env)
|
||||||
|
- `-base-url` — upstream inference engine base URL for BYOK mode (`OPENAI_BASE_URL` / `QWEN_BASE_URL` env)
|
||||||
|
- `-socks` / `-proxy` / `-socks5` — SOCKS5 proxy URL, e.g. `socks5://127.0.0.1:1080` (`ALL_PROXY` env)
|
||||||
|
- `-user-agent` / `-ua` — custom User-Agent sent to upstream
|
||||||
|
|
||||||
|
Endpoints served:
|
||||||
|
|
||||||
|
- `GET /models` and `GET /v1/models`
|
||||||
|
- `POST /chat/completions` and `POST /v1/chat/completions`
|
||||||
|
|
||||||
|
Example request with curl:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl http://127.0.0.1:8080/v1/models
|
curl http://localhost:8080/v1/chat/completions \
|
||||||
```
|
|
||||||
|
|
||||||
### 3. Non-Streaming Chat Completion
|
|
||||||
|
|
||||||
```bash
|
|
||||||
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
|
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{
|
-d '{"model":"qwen3.8-flash-next","messages":[{"role":"user","content":"Explain QSA micro-blocks in one sentence."}],"stream":false}'
|
||||||
"model": "qwen3.8-flash-next",
|
|
||||||
"messages": [
|
|
||||||
{"role": "user", "content": "Explain QSA micro-blocks in one sentence."}
|
|
||||||
],
|
|
||||||
"stream": false
|
|
||||||
}'
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4. Streaming Chat Completion (Real-Time SSE)
|
Streaming example:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -N -X POST http://127.0.0.1:8080/v1/chat/completions \
|
curl -N http://localhost:8080/v1/chat/completions \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{
|
-d '{"model":"qwen-flash","messages":[{"role":"user","content":"Count from 1 to 5."}],"stream":true}'
|
||||||
"model": "qwen-flash",
|
|
||||||
"messages": [
|
|
||||||
{"role": "user", "content": "Write a quick Python countdown loop."}
|
|
||||||
],
|
|
||||||
"stream": true
|
|
||||||
}'
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 5. Instruct Mode (Disable Thinking)
|
Instruct mode example (disables thinking):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
|
curl http://localhost:8080/v1/chat/completions \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{
|
-d '{"model":"qwen-flash","messages":[{"role":"user","content":"Hello!"}],"reasoning_effort":"none"}'
|
||||||
"model": "qwen-flash",
|
|
||||||
"messages": [
|
|
||||||
{"role": "user", "content": "Hello!"}
|
|
||||||
],
|
|
||||||
"reasoning_effort": "none"
|
|
||||||
}'
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 6. Python OpenAI SDK Integration
|
Python OpenAI SDK integration:
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from openai import OpenAI
|
from openai import OpenAI
|
||||||
|
|
||||||
client = OpenAI(
|
client = OpenAI(
|
||||||
base_url="http://127.0.0.1:8080/v1",
|
base_url="http://localhost:8080/v1",
|
||||||
api_key="sk-dummy"
|
api_key="sk-dummy"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -169,3 +132,7 @@ for chunk in stream:
|
|||||||
print(delta.content, end="", flush=True)
|
print(delta.content, end="", flush=True)
|
||||||
print()
|
print()
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Credits
|
||||||
|
|
||||||
|
Created by Luxferre in 2026, released into the public domain with no warranties.
|
||||||
|
|||||||
Reference in New Issue
Block a user