# Qflash: OpenAI-compatible gateway for Qwen3.8 models ## About Qflash is a standalone, single-binary gateway that exposes Qwen3.8 model spaces on Hugging Face through a standard OpenAI-compatible API. It automatically detects and supports multiple upstream protocols: - direct OpenAI `/v1/chat/completions` endpoints (such as `https://wanyamaelis-qwen3-8-27b.hf.space` and `https://apathy-exe-qwen3-8-flash-next.hf.space`, running persistent `llama-server` instances on CPU with zero quota limits) - gradio `/respond` endpoints (supported for custom spaces such as `https://microhero-qwen3-8-27b-uncensored-chat.hf.space`) - gradio `/chat_response` endpoints (such as `https://halvo78-qwen3-8-flash-next-playground.hf.space`) It translates standard `/v1/chat/completions` and `/v1/models` requests and Server-Sent Events (SSE) stream protocols into upstream formats, allowing any standard OpenAI-compatible client, agent, or IDE to interface with Qwen3.8 models without modification. ## Features - openai-compatible chat completions (streaming and non-streaming) - automatic upstream failover across configured non-ZeroGPU endpoints - smart endpoint cooldown (5 minutes on quota exhaustion, 30 seconds on network errors) - automatic upstream endpoint detection (`respond`, `openai`, `chat_response`) - deep reasoning extraction with thinking trace passthrough (`` tags and blockquotes mapped to `reasoning_content`) - stateful streaming tool call interception (`StreamToolCallFilter`) with zero XML or JSON leakage into `delta.content` - real-time token streaming with incremental SSE delivery - support for instruct mode via standard `reasoning_effort: "none"` - zero-dependency SOCKS5 proxy client (RFC 1928 / RFC 1929) with domain resolution (`socks5h://`), IPv4, IPv6, and auth - bring your own key (BYOK) pass-through support via `Authorization: Bearer` or CLI flags - zero-gpu quota authentication via free Hugging Face personal tokens (`HF_TOKEN`) - fibonacci backoff retry on transient upstream errors - zero external dependencies (Go standard library only) ## Supported upstream spaces | Space | Model | Hardware | Protocol | Notes | |---|---|---|---|---| | `Wanyamaelis/Qwen3.8-27B` *(primary)* | Qwen3.8-27B (MTP) | CPU basic (8 vCPU) | Native OpenAI `/v1` | Default endpoint, non-ZeroGPU, no quota limits, fast speculative decoding (~5s) | | `apathy-exe/Qwen3.8-27B` *(fallback)* | Qwen3.8-27B (MTP) | CPU (OpenMP/AVX512) | Native OpenAI `/v1` | Non-ZeroGPU, no quota limits, speculative decoding | | `apathy-exe/Qwen3.8-Flash-Next` *(fallback)* | Qwen3.8-Flash-Next (~177B) | CPU (OpenMP/AVX512) | Native OpenAI `/v1` | Non-ZeroGPU, no quota limits, full 177B Flash-Next model | | `MicroHERO/qwen3.8-27b-uncensored-chat` | Qwen3.8-27B Uncensored | ZeroGPU (A10G) | Gradio `/respond` | Fast live GPU inference, can be configured via custom `-endpoints` | | `Halvo78/qwen3-8-flash-next-playground` | Qwen3.8-Flash-Next | CPU basic | Gradio `/chat_response` | Sandbox client; requires BYOK API key/base URL | ### Auto-failover and quota handling Default endpoints run on persistent non-ZeroGPU hardware with no daily runs limit. When custom ZeroGPU spaces are configured, anonymous requests share a small pool per IP address. When runs limit is reached, upstream returns a quota error (`429` or `ZeroGPU runs limit`). Qflash handles upstream reliability seamlessly: - with auto-failover enabled (default), when an endpoint encounters an error or timeout, it is placed on cooldown and the gateway automatically fails over to the next configured endpoint - transient network errors trigger a 30-second cooldown, while quota limits trigger a 5-minute cooldown - optional ZeroGPU or private spaces can be authenticated using a Hugging Face personal access token (`https://huggingface.co/settings/tokens`) via the `HF_TOKEN` environment variable, the `-hf-token` CLI flag, or the `Authorization: Bearer hf_...` header ## Installation You need Go 1.22 or newer (tested on Go 1.26). Install the latest release straight from the Git repository with `go install`: ```bash go install code.luxferre.top/luxferre/qflash@latest ``` This fetches the module from `https://code.luxferre.top/luxferre/qflash.git` and places the `qflash` binary in `$(go env GOPATH)/bin`. Make sure that directory is on your `PATH`. *(Note: If installing right after a new commit has been pushed, bypass any proxy cache with `GOPROXY=direct go install code.luxferre.top/luxferre/qflash@latest`).* If you prefer to build from a local checkout instead: ```bash git clone https://code.luxferre.top/luxferre/qflash.git cd qflash go install . ``` Alternatively, build directly from source using `make`: ```bash make qflash ``` This produces the `bin/qflash` binary for your platform. ## Models served The gateway advertises the following models under `/v1/models`: | Model ID | Target model | Description | |---|---|---| | `Qwen/Qwen3.8-27B` | `Qwen/Qwen3.8-27B` | Default primary live model | | `Qwen/Qwen3.8-27B-Uncensored` | `Qwen/Qwen3.8-27B-Uncensored` | Uncensored model identifier | | `Qwen/Qwen3.8-Flash-Next` | `Qwen/Qwen3.8-Flash-Next` | Flash-Next model identifier | | `qwen3.8-27b` | `Qwen/Qwen3.8-27B` | Standard lowercase alias | | `qwen3.8-27b-uncensored` | `Qwen/Qwen3.8-27B-Uncensored` | Uncensored lowercase alias | | `qwen3.8-flash-next` | `Qwen/Qwen3.8-Flash-Next` | Flash-Next lowercase alias | | `qwen-flash-next` | `Qwen/Qwen3.8-Flash-Next` | Shorthand alias | | `qwen-flash` | `Qwen/Qwen3.8-Flash-Next` | Quick convenience alias | | `qwen` | Default model | Generic shorthand alias | Any unlisted custom model name requested by the client is passed through directly. ## Usage Run the gateway with default auto-failover endpoints: ```bash qflash ``` By default, this listens on `http://127.0.0.1:8080` with failover configured across persistent non-ZeroGPU endpoints (`wanyamaelis` and `apathy-exe`). Available flags: - `-port` — tcp port to listen on (default `8080`) - `-endpoints` / `-endpoint` — comma-separated upstream space or OpenAI URLs (default list of 3 non-ZeroGPU endpoints, `QFLASH_ENDPOINTS` / `QFLASH_ENDPOINT` env) - `-failover` / `-auto-failover` — enable automatic failover across endpoints on quota exhaustion or error (default `true`, `QFLASH_FAILOVER` env) - `-mode` — upstream protocol mode: `auto`, `respond`, `chat_response`, `openai` (default `auto`, `QFLASH_MODE` env) - `-model` — exposed model name (default `Qwen/Qwen3.8-27B`, `QFLASH_MODEL` env) - `-thinking` / `-enable-thinking` — enable chain-of-thought reasoning by default (default `true`) - `-hf-token` / `-token` — Hugging Face API token for ZeroGPU quota or private spaces (`HF_TOKEN` env) - `-api-key` — upstream inference engine API key for BYOK mode (`OPENAI_API_KEY` / `QWEN_API_KEY` env) - `-base-url` — upstream inference engine base URL for BYOK mode (`OPENAI_BASE_URL` / `QWEN_BASE_URL` env) - `-socks` / `-proxy` / `-socks5` — SOCKS5 proxy URL, e.g. `socks5://127.0.0.1:1080` (`ALL_PROXY` env) - `-user-agent` / `-ua` — custom User-Agent sent to upstream Endpoints served: - `GET /models` and `GET /v1/models` - `POST /chat/completions` and `POST /v1/chat/completions` Example request with curl: ```bash curl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"qwen","messages":[{"role":"user","content":"Explain quantum superposition in one sentence."}],"stream":false}' ``` Streaming example: ```bash curl -N http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"qwen","messages":[{"role":"user","content":"Count from 1 to 5."}],"stream":true}' ``` Instruct mode example (disables thinking): ```bash curl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"qwen","messages":[{"role":"user","content":"Hello!"}],"reasoning_effort":"none"}' ``` Using custom upstream spaces or a single endpoint: ```bash # Point to a single CPU llama-server without failover qflash -endpoint https://wanyamaelis-qwen3-8-27b.hf.space -failover=false # Point to Halvo78 playground with your own API credentials qflash -endpoint https://halvo78-qwen3-8-flash-next-playground.hf.space -api-key "sk-..." -base-url "https://api.openai.com/v1" ``` Python OpenAI SDK integration: ```python from openai import OpenAI client = OpenAI( base_url="http://localhost:8080/v1", api_key="sk-dummy" ) stream = client.chat.completions.create( model="qwen", messages=[ {"role": "user", "content": "Prove that the sum of the first n odd numbers is n^2."} ], stream=True ) for chunk in stream: delta = chunk.choices[0].delta if hasattr(delta, "reasoning_content") and delta.reasoning_content: print(f"[THINK] {delta.reasoning_content}", end="", flush=True) if delta.content: print(delta.content, end="", flush=True) print() ``` ## Credits Created by Luxferre in 2026, released into the public domain with no warranties.