7623220e17f145dae1f37d955190db1f153bb171
qflash: OpenAI Proxy Gateway for Qwen3.8-Flash-Next
Standalone, performant, zero-dependency Go OpenAI proxy gateway for the Qwen3.8-Flash-Next Hugging Face Gradio space (https://halvo78-qwen3-8-flash-next-playground.hf.space).
Created by Luxferre in 2026, released into the public domain.
Features
- Zero External Dependencies: Built entirely with Go standard library packages (
net/http,encoding/json,bufio, etc.). - OpenAI-Compatible API: Implements standard
/v1/chat/completions(streaming & non-streaming) and/v1/models. - Real-Time Token Streaming: Streams SSE chunks with incremental token delivery directly to clients.
- Deep Reasoning Separation:
- Automatically isolates thinking traces from both standard
<think>...</think>tags and the playground's blockquote thinking blocks (> 💭 **Thinking Process...**). - Emits pure thought traces to
delta.reasoning_content(streaming) andmessage.reasoning_content(non-streaming). - Keeps
delta.contentandmessage.contentclean.
- Automatically isolates thinking traces from both standard
- Stateful Streaming Tool Call Interception:
- Injects tool schemas into system instructions.
- Intercepts
<tool_call>blocks in real time viaStreamToolCallFilterwithout leaking raw XML or JSON intodelta.content. - Emits structured
delta.tool_callschunks and setsfinish_reason: "tool_calls".
- Reasoning Effort Control: Respects standard
reasoning_effort: "none"to switch dynamically into high-speed Instruct Mode. - Zero-Dependency SOCKS5 Proxy Client:
- RFC 1928 and RFC 1929 compliant client with domain resolution (
socks5h://), IPv4, IPv6, and authentication. - Wireable via
-socksCLI flag orALL_PROXY/SOCKS5_PROXYenvironment variables.
- RFC 1928 and RFC 1929 compliant client with domain resolution (
- Bring Your Own Key (BYOK) Pass-through:
- Passes client API keys or custom base URLs directly to upstream inference engines when provided.
Architecture & Model Aliases
The gateway serves the following models under /v1/models:
| Model ID | Target Model | Description |
|---|---|---|
Qwen/Qwen3.8-Flash-Next |
Qwen/Qwen3.8-Flash-Next |
Primary playground model (125B MoE, 6B activated) |
qwen3.8-flash-next |
Qwen/Qwen3.8-Flash-Next |
Standard lowercase alias |
qwen-flash-next |
Qwen/Qwen3.8-Flash-Next |
Shorthand alias |
qwen-flash |
Qwen/Qwen3.8-Flash-Next |
Quick convenience alias |
Any unlisted custom model name requested by the client is passed through directly.
Build Instructions
Build binary with Go:
make qflash
Or run test suite:
make test
The resulting binary will be placed at bin/qflash.
Configuration Flags & Environment Variables
| Flag | Shorthand | Environment Variable | Default | Description |
|---|---|---|---|---|
-port |
PORT |
8080 |
Port to bind the HTTP server | |
-endpoint |
https://halvo78-qwen3-8-flash-next-playground.hf.space |
Upstream Gradio space base URL | ||
-model |
Qwen/Qwen3.8-Flash-Next |
Default model ID | ||
-thinking |
-enable-thinking |
true |
Enable chain-of-thought reasoning by default | |
-hf-token |
HF_TOKEN |
"" |
Hugging Face user access token | |
-api-key |
OPENAI_API_KEY / QWEN_API_KEY |
"" |
Upstream inference engine API key | |
-base-url |
OPENAI_BASE_URL / QWEN_BASE_URL |
"" |
Upstream inference engine base URL | |
-user-agent |
-ua |
USER_AGENT |
Firefox 153 on Linux | Custom User-Agent header |
-socks |
-proxy, -socks5 |
ALL_PROXY, SOCKS5_PROXY |
"" |
SOCKS5 proxy URL (socks5://127.0.0.1:1080) |
Usage Examples
1. Launch Gateway
./bin/qflash -port 8080
2. List Models
curl http://127.0.0.1:8080/v1/models
3. Non-Streaming Chat Completion
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-flash-next",
"messages": [
{"role": "user", "content": "Explain QSA micro-blocks in one sentence."}
],
"stream": false
}'
4. Streaming Chat Completion (Real-Time SSE)
curl -N -X POST http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-flash",
"messages": [
{"role": "user", "content": "Write a quick Python countdown loop."}
],
"stream": true
}'
5. Instruct Mode (Disable Thinking)
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-flash",
"messages": [
{"role": "user", "content": "Hello!"}
],
"reasoning_effort": "none"
}'
6. Python OpenAI SDK Integration
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8080/v1",
api_key="sk-dummy"
)
stream = client.chat.completions.create(
model="qwen3.8-flash-next",
messages=[
{"role": "user", "content": "Prove that the sum of the first n odd numbers is n^2."}
],
stream=True
)
for chunk in stream:
delta = chunk.choices[0].delta
if hasattr(delta, "reasoning_content") and delta.reasoning_content:
print(f"[THINK] {delta.reasoning_content}", end="", flush=True)
if delta.content:
print(delta.content, end="", flush=True)
print()
Languages
Go
99.7%
Makefile
0.3%