7.3 KiB
7.3 KiB
gr2gw: Universal Gradio to OpenAI LLM gateway
A zero-dependency, high-performance Go proxy server that introspects any Gradio chat space (such as Hugging Face Spaces or custom deployments) and exposes a standards-compliant OpenAI /v1/chat/completions and /v1/models HTTP API.
Default demo space: https://ghost2513-openai-gpt-oss-120b.hf.space
Features
- Zero external dependencies: Pure Go standard library (
net/http,encoding/json,bufio, etc.). - Automatic space introspection: Dynamically queries
/gradio_api/info,/config, and Hugging Face space metadata to discover models, endpoints, and input parameter mappings. - Native Tencent Hunyuan 3 (
tencent-hy3) support:- Full native zero-degradation handling for official spaces like
https://tencent-hy3.hf.space. - Maps
functions_json_strnatively without polluting the system prompt. - Preserves multi-turn reasoning content and tool call history in standard OpenAI message schemas.
- Maps
reasoning_effort(no_think,low,high) directly tothink_level.
- Full native zero-degradation handling for official spaces like
- Universal multi-turn handling:
- Automatically formats conversation history into structured inputs when the space supports them.
- Transparently composes multi-turn dialogue (
System,User,Assistant) into single prompt inputs when the space only accepts a single message textbox. - Automatically pads hidden/State inputs (e.g. Gradio State components) to prevent backend argument count mismatches.
- Real-time streaming & accumulation filter:
- Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames
[[content, reasoning, tool_calls, history]]). - Emits standards-compliant
chat.completion.chunkSSE events in real time.
- Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames
- Thinking & reasoning token separation:
- Streams native reasoning chunks as
delta.reasoning_contentin real time. - Detects
<think>...</think>tags in real time as fallback for standard spaces. - Separates reasoning into
delta.reasoning_content(streaming) andmessage.reasoning_content(non-streaming). - Keeps
contentclean without tag leakage.
- Streams native reasoning chunks as
- Full tool calling & function interception:
- Formats schemas into native
functions_json_str(Hy3) or system prompts (standard spaces). StreamToolCallFilter: Stateful sliding-window filter that prevents<tool_call>tags from leaking intodelta.content. Emits structured OpenAIdelta.tool_callschunks and setsfinish_reason: "tool_calls".- Seamlessly maintains multi-turn context when tool results are submitted back via
role: "tool".
- Formats schemas into native
- Built-in SOCKS5 proxy client:
- Full RFC 1928 / RFC 1929 implementation with domain resolution (
socks5h://), IPv4, IPv6, and username/password auth.
- Full RFC 1928 / RFC 1929 implementation with domain resolution (
- Dynamic space override:
- Switch the target Gradio space on-the-fly per request using the
X-Gradio-SpaceorX-Space-URLHTTP headers.
- Switch the target Gradio space on-the-fly per request using the
- Fibonacci retry engine:
- Resilient backoff retry mechanism (1s, 1s, 2s, 3s, 5s) for transient network hiccups.
Installation
go install code.luxferre.top/luxferre/gr2gw@latest
Build
make build
Binary will be compiled to bin/gr2gw.
To run tests:
make test
Usage
Quick start
Run with the default space (https://ghost2513-openai-gpt-oss-120b.hf.space):
./bin/gr2gw -port 8080
Target any other Gradio space:
./bin/gr2gw -space https://ericsqin-hy3.hf.space -port 8080
With SOCKS5 proxy:
./bin/gr2gw -space https://ghost2513-openai-gpt-oss-120b.hf.space -socks socks5://127.0.0.1:1080
CLI flags
| Flag | Default | Description |
|---|---|---|
-space, -url |
https://ghost2513-openai-gpt-oss-120b.hf.space |
Target Gradio space URL |
-port |
8080 |
Port to listen on |
-host |
0.0.0.0 |
Host interface to bind to |
-socks, -proxy, -socks5 |
"" |
SOCKS5 proxy URL (socks5://user:pass@host:port) |
-user-agent, -ua |
Firefox string | Custom User-Agent header |
-timeout |
300 |
Upstream request timeout in seconds |
Environment variables
GRADIO_SPACE_URL: default Gradio space URL fallback.ALL_PROXY,SOCKS5_PROXY,SOCKS_PROXY: default SOCKS5 proxy URL fallback.
API examples
List models
curl http://localhost:8080/v1/models
Response:
{
"object": "list",
"data": [
{
"id": "openai/gpt-oss-120b",
"object": "model",
"created": 1788756307,
"owned_by": "gradio"
},
{
"id": "gpt-oss-120b",
"object": "model",
"created": 1788756307,
"owned_by": "gradio"
}
]
}
Chat completions (non-streaming)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b",
"messages": [
{"role": "user", "content": "What is the capital of France?"}
]
}'
Response:
{
"id": "chatcmpl-16425f9d-c350-47a1-9a6d-e9ce10871545",
"object": "chat.completion",
"created": 1788756310,
"model": "openai/gpt-oss-120b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Paris is the capital of France."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
Chat completions (streaming)
curl -N http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b",
"messages": [
{"role": "user", "content": "Count from 1 to 5."}
],
"stream": true
}'
Tool calling
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b",
"messages": [
{"role": "user", "content": "What is the weather in Tokyo?"}
],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather in a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}]
}'
Response:
{
"id": "chatcmpl-32727c62-ef2a-4866-855b-f1c7ec2b8023",
"object": "chat.completion",
"created": 1788756322,
"model": "openai/gpt-oss-120b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_3d4c016a",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\":\"Tokyo\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
Dynamic target space override
Override the target space per request without restarting the server:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-Gradio-Space: https://ericsqin-hy3.hf.space" \
-d '{
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
Credits
Created by Luxferre in 2026, released into the public domain with no warranties.