15 KiB
gr2gw: universal Gradio to OpenAI LLM gateway
A zero-dependency, high-performance Go proxy server that introspects any Gradio chat space (such as Hugging Face Spaces or custom deployments) and exposes a standards-compliant OpenAI /v1/chat/completions and /v1/models HTTP API.
Default demo space: https://tencent-hy3.hf.space
Features
- Zero external dependencies: pure Go standard library (
net/http,encoding/json,bufio, etc.). - Universal heuristic discovery engine:
- Automatically interrogates
/gradio_api/info,/config, and Hugging Face metadata without endpoint-specific hardcoding. - Detects Gradio runtime versions across v3, v4, v5, and v6.
- Classifies application architecture into
ChatInterface,Blocks (Chat),Blocks (Multimodal Chat),Interface, andGeneric. - Disambiguation scoring engine evaluates candidate endpoints, filtering out UI resets, clears, retries, likes, and utility triggers to pinpoint primary conversational completion functions.
- Correlates semantic parameter names from
/gradio_api/infowith component IDs from/configto reconstruct parameter mappings even when component labels are obfuscated.
- Automatically interrogates
- Dual protocol support with auto-fallback:
- Supports modern Gradio 4/5/6
/callSSE protocol with persistent session hashes. - Supports legacy Gradio 3
/run/predictand/api/predictdirect execution protocol. - Instant zero-latency fallback from
/callto/run/predictupon HTTP 404 or 405 status codes.
- Supports modern Gradio 4/5/6
- Tool calling support classification:
- Automatically identifies tool calling mechanisms:
native_slot,prompt_augmented_system,prompt_augmented_first_turn, orprompt_augmented_single_prompt. - Upstream tool call recovery intercepts
tool_use_failederrors and extracts function names and JSON arguments. - Real-time sliding-window
<tool_call>tag interceptor emits structured OpenAI tool call chunks.
- Automatically identifies tool calling mechanisms:
- Native Tencent Hunyuan 3 (
tencent-hy3) support:- Full native zero-degradation handling for official spaces like
https://tencent-hy3.hf.space. - Maps
functions_json_strnatively without polluting the system prompt. - Preserves multi-turn reasoning content and tool call history in standard OpenAI message schemas.
- Maps
reasoning_effort(no_think,low,high) directly tothink_level.
- Full native zero-degradation handling for official spaces like
- Universal multi-turn handling:
- Automatically formats conversation history into structured inputs when the space supports them.
- Transparently composes multi-turn dialogue (
System,User,Assistant) into single prompt inputs when the space only accepts a single message textbox. - Automatically pads hidden/State inputs (e.g. Gradio State components) to prevent backend argument count mismatches.
- Automatically handles multimodal textbox inputs (
MultimodalDatawith{text, files}) and space component defaults (radios, sliders, checkboxes).
- Real-time streaming & accumulation filter:
- Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames
[[content, reasoning, tool_calls, history]]). - Emits standards-compliant
chat.completion.chunkSSE events in real time.
- Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames
- Thinking & reasoning token separation:
- Streams native reasoning chunks as
delta.reasoning_contentin real time. - Detects
<think>...</think>tags in real time as fallback for standard spaces. - Separates reasoning into
delta.reasoning_content(streaming) andmessage.reasoning_content(non-streaming). - Keeps
contentclean without tag leakage.
- Streams native reasoning chunks as
- Built-in SOCKS5 proxy client:
- Full RFC 1928 / RFC 1929 implementation with domain resolution (
socks5h://), IPv4, IPv6, and username/password auth.
- Full RFC 1928 / RFC 1929 implementation with domain resolution (
- Dynamic space override:
- Switch the target Gradio space on-the-fly per request using the
X-Gradio-SpaceorX-Space-URLHTTP headers.
- Switch the target Gradio space on-the-fly per request using the
- Fibonacci retry engine:
- Resilient backoff retry mechanism (1s, 1s, 2s, 3s, 5s) for transient network hiccups with immediate break on 404/405 errors.
Heuristic discovery engine
The gateway implements an autonomous heuristic engine that discovers and configures the optimal completion path for any target Gradio space at startup.
Space introspection and version detection
Upon initialization, gr2gw inspects the space metadata:
- Queries
/gradio_api/infoand/configendpoints. - Extracts the Gradio runtime version (v3, v4, v5, or v6).
- Detects API routing prefixes (e.g.
/gradio_apion modern versions, or root on Gradio 3).
UI flavor classification
The engine classifies the application structure into architectural flavors:
ChatInterface: Standard Gradio chat interfaces equipped with chatbot, textbox, and optional additional inputs.Blocks (Chat): Customgr.Blockslayouts containing conversational components.Blocks (Multimodal Chat): Blocks architectures featuringMultimodalTextboxcomponents that accept{text, files}JSON payloads.Interface: Classic input-outputgr.Interfaceinstances.Generic: Spaces with custom or unclassified component topologies.
Candidate endpoint scoring
Gradio spaces frequently expose dozens of internal endpoints for UI actions (e.g. clearing text, retrying responses, voting/liking, adjusting sliders). The scoring algorithm identifies the true conversational endpoint by:
- Penalizing non-conversational triggers (e.g.
-600for clear/reset/undo/retry/like endpoints). - Rewarding chat semantics (
+150for/chat,/predict,/generate,/respond). - Rewarding message inputs (
+120forTextboxorMultimodalTextbox). - Rewarding chat history slots (
+80forChatbotorStatecomponents). - Rewarding generator and streaming dependencies (
+50).
Dual protocol execution and auto-fallback
callprotocol: Modern Gradio 4/5/6 execution viaPOST /call/{endpoint}returning an event ID, followed byGET /call/{endpoint}/{event_id}SSE streaming.predictprotocol: Gradio 3 and legacy execution via directPOST /run/predictorPOST /api/predict.- Runtime failover: If a space returns HTTP 404 or 405 when calling the modern protocol, the gateway breaks immediately from the retry loop and falls back to
/run/predict.
Tool calling support modes
The gateway evaluates available input components to determine how tool schemas and function calls should be delivered:
native_slot: The space provides a dedicated parameter slot for tool definitions (e.g.functions_json_stron Hunyuan 3). Function definitions are passed cleanly without prompt alteration.prompt_augmented_system: The space provides a separatesystem_promptinput slot. Tool definitions and invocation schemas are injected directly into the system prompt.prompt_augmented_first_turn: The space accepts chat history pairs but lacks a dedicated system prompt slot. Tool definitions are prepended to the user prompt on the first dialogue turn.prompt_augmented_single_prompt: The space accepts only a single textbox input. Full multi-turn dialogue, tool definitions, and system guidance are synthesized into a single cohesive prompt.
Startup resolution diagnostics
Whenever the gateway starts or inspects a new space, it prints the complete resolution picture:
================================================================================
Gradio Space Resolution Picture
--------------------------------------------------------------------------------
Space URL: https://tencent-hy3.hf.space
Title: Hunyuan 3 Chat
Gradio Version: 5.29.0
UI Flavor: Blocks (Chat)
Protocol: call
API Prefix: /gradio_api
Resolved Endpoint: /chat_fn
Function Index: 1
Primary Model: hy3
Exposed Models: hy3, hunyuan3, tencent/Hy3
History Format: tuples
Tool Call Support: native_slot
Total Input Slots: 9
Input Slot Mappings:
[0] Component ID 1 textbox (label="message") -> message
[1] Component ID 2 textbox (label="system") -> system_prompt
[2] Component ID 3 chatbot (label="chatbot") -> history
[3] Component ID 4 radio (label="think_level") -> think_level
[4] Component ID 5 slider (label="temperature") -> temperature
[5] Component ID 6 slider (label="max_tokens") -> max_tokens
[6] Component ID 7 slider (label="top_p") -> top_p
[7] Component ID 8 state (label="preserved") -> preserved_thinking
[8] Component ID 9 textbox (label="functions") -> functions_json_str
================================================================================
Gateway status endpoint
Making a GET request to / returns a JSON summary of the running gateway and the discovered space profile:
curl http://localhost:8080/
Response:
{
"name": "gr2gw",
"status": "ready",
"space_url": "https://tencent-hy3.hf.space",
"gradio_version": "5.29.0",
"flavor": "Blocks (Chat)",
"protocol": "call",
"tool_call_mode": "native_slot",
"models": ["hy3", "hunyuan3", "tencent/Hy3"]
}
Installation
go install code.luxferre.top/luxferre/gr2gw@latest
Build
make build
Binary will be compiled to bin/gr2gw.
To run tests:
make test
Usage
Quick start
Run with the default space (https://tencent-hy3.hf.space):
./bin/gr2gw -port 8080
Target any other Gradio space:
./bin/gr2gw -space https://lucasmarchettidelima-digital-twin.hf.space -port 8080
With SOCKS5 proxy:
./bin/gr2gw -space https://tencent-hy3.hf.space -socks socks5://127.0.0.1:1080
CLI flags
| Flag | Default | Description |
|---|---|---|
-space, -url, -endpoint |
https://tencent-hy3.hf.space |
Target Gradio space URL |
-model |
"" |
Exposed model name override (default: auto-detected) |
-port |
8080 |
Port to listen on |
-host |
0.0.0.0 |
Host interface to bind to |
-socks, -proxy, -socks5 |
"" |
SOCKS5 proxy URL (socks5://user:pass@host:port) |
-user-agent, -ua |
Firefox string | Custom User-Agent header |
-timeout |
300 |
Upstream request timeout in seconds |
Environment variables
GRADIO_SPACE_URL: default Gradio space URL fallback.ALL_PROXY,SOCKS5_PROXY,SOCKS_PROXY: default SOCKS5 proxy URL fallback.
API examples
List models
curl http://localhost:8080/v1/models
Response:
{
"object": "list",
"data": [
{
"id": "hy3",
"object": "model",
"created": 1788756307,
"owned_by": "gradio"
},
{
"id": "hunyuan3",
"object": "model",
"created": 1788756307,
"owned_by": "gradio"
},
{
"id": "tencent/Hy3",
"object": "model",
"created": 1788756307,
"owned_by": "gradio"
}
]
}
Chat completions (non-streaming)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "hy3",
"messages": [
{"role": "user", "content": "What is the capital of France?"}
]
}'
Response:
{
"id": "chatcmpl-16425f9d-c350-47a1-9a6d-e9ce10871545",
"object": "chat.completion",
"created": 1788756310,
"model": "hy3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The capital of France is Paris."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
Chat completions (streaming)
curl -N http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "hy3",
"messages": [
{"role": "user", "content": "Count from 1 to 5."}
],
"stream": true
}'
Tool calling (turn 1)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "hy3",
"messages": [
{"role": "user", "content": "What is the weather in Tokyo?"}
],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather in a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}]
}'
Response:
{
"id": "chatcmpl-32727c62-ef2a-4866-855b-f1c7ec2b8023",
"object": "chat.completion",
"created": 1788756322,
"model": "hy3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_3d4c016a",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\":\"Tokyo\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
Tool response submission (turn 2)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "hy3",
"messages": [
{"role": "user", "content": "What is the weather in Tokyo?"},
{
"role": "assistant",
"tool_calls": [
{
"id": "call_3d4c016a",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"location\":\"Tokyo\"}"
}
}
]
},
{
"role": "tool",
"tool_call_id": "call_3d4c016a",
"content": "{\"temperature\": 20, \"condition\": \"sunny\"}"
}
],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather in a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}]
}'
Response:
{
"id": "chatcmpl-4903ba12-f12b-4cd3-a801-7290bc91a421",
"object": "chat.completion",
"created": 1788756335,
"model": "hy3",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The current weather in Tokyo is sunny with a temperature of 20 °C."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
Dynamic target space override
Override the target space per request without restarting the server:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-H "X-Gradio-Space: https://ericsqin-hy3.hf.space" \
-d '{
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
Credits
Created by Luxferre in 2026, released into the public domain with no warranties.