qflash: OpenAI Proxy Gateway for Qwen3.8-Flash-Next

Standalone, performant, zero-dependency Go OpenAI proxy gateway for the Qwen3.8-Flash-Next Hugging Face Gradio space (https://halvo78-qwen3-8-flash-next-playground.hf.space).

Created by Luxferre in 2026, released into the public domain.


Features

  • Zero External Dependencies: Built entirely with Go standard library packages (net/http, encoding/json, bufio, etc.).
  • OpenAI-Compatible API: Implements standard /v1/chat/completions (streaming & non-streaming) and /v1/models.
  • Real-Time Token Streaming: Streams SSE chunks with incremental token delivery directly to clients.
  • Deep Reasoning Separation:
    • Automatically isolates thinking traces from both standard <think>...</think> tags and the playground's blockquote thinking blocks (> 💭 **Thinking Process...**).
    • Emits pure thought traces to delta.reasoning_content (streaming) and message.reasoning_content (non-streaming).
    • Keeps delta.content and message.content clean.
  • Stateful Streaming Tool Call Interception:
    • Injects tool schemas into system instructions.
    • Intercepts <tool_call> blocks in real time via StreamToolCallFilter without leaking raw XML or JSON into delta.content.
    • Emits structured delta.tool_calls chunks and sets finish_reason: "tool_calls".
  • Reasoning Effort Control: Respects standard reasoning_effort: "none" to switch dynamically into high-speed Instruct Mode.
  • Zero-Dependency SOCKS5 Proxy Client:
    • RFC 1928 and RFC 1929 compliant client with domain resolution (socks5h://), IPv4, IPv6, and authentication.
    • Wireable via -socks CLI flag or ALL_PROXY / SOCKS5_PROXY environment variables.
  • Bring Your Own Key (BYOK) Pass-through:
    • Passes client API keys or custom base URLs directly to upstream inference engines when provided.

Architecture & Model Aliases

The gateway serves the following models under /v1/models:

Model ID Target Model Description
Qwen/Qwen3.8-Flash-Next Qwen/Qwen3.8-Flash-Next Primary playground model (125B MoE, 6B activated)
qwen3.8-flash-next Qwen/Qwen3.8-Flash-Next Standard lowercase alias
qwen-flash-next Qwen/Qwen3.8-Flash-Next Shorthand alias
qwen-flash Qwen/Qwen3.8-Flash-Next Quick convenience alias

Any unlisted custom model name requested by the client is passed through directly.


Build Instructions

Build binary with Go:

make qflash

Or run test suite:

make test

The resulting binary will be placed at bin/qflash.


Configuration Flags & Environment Variables

Flag Shorthand Environment Variable Default Description
-port PORT 8080 Port to bind the HTTP server
-endpoint https://halvo78-qwen3-8-flash-next-playground.hf.space Upstream Gradio space base URL
-model Qwen/Qwen3.8-Flash-Next Default model ID
-thinking -enable-thinking true Enable chain-of-thought reasoning by default
-hf-token HF_TOKEN "" Hugging Face user access token
-api-key OPENAI_API_KEY / QWEN_API_KEY "" Upstream inference engine API key
-base-url OPENAI_BASE_URL / QWEN_BASE_URL "" Upstream inference engine base URL
-user-agent -ua USER_AGENT Firefox 153 on Linux Custom User-Agent header
-socks -proxy, -socks5 ALL_PROXY, SOCKS5_PROXY "" SOCKS5 proxy URL (socks5://127.0.0.1:1080)

Usage Examples

1. Launch Gateway

./bin/qflash -port 8080

2. List Models

curl http://127.0.0.1:8080/v1/models

3. Non-Streaming Chat Completion

curl -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-flash-next",
    "messages": [
      {"role": "user", "content": "Explain QSA micro-blocks in one sentence."}
    ],
    "stream": false
  }'

4. Streaming Chat Completion (Real-Time SSE)

curl -N -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-flash",
    "messages": [
      {"role": "user", "content": "Write a quick Python countdown loop."}
    ],
    "stream": true
  }'

5. Instruct Mode (Disable Thinking)

curl -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-flash",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ],
    "reasoning_effort": "none"
  }'

6. Python OpenAI SDK Integration

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key="sk-dummy"
)

stream = client.chat.completions.create(
    model="qwen3.8-flash-next",
    messages=[
        {"role": "user", "content": "Prove that the sum of the first n odd numbers is n^2."}
    ],
    stream=True
)

for chunk in stream:
    delta = chunk.choices[0].delta
    if hasattr(delta, "reasoning_content") and delta.reasoning_content:
        print(f"[THINK] {delta.reasoning_content}", end="", flush=True)
    if delta.content:
        print(delta.content, end="", flush=True)
print()
S
Description
Free Qwen 3.8 Flash Next LLM gateway
Readme
5.6 MiB
Languages
Go 99.7%
Makefile 0.3%