Files
gr2gw/README.md
T

15 KiB

gr2gw: universal Gradio to OpenAI LLM gateway

A zero-dependency, high-performance Go proxy server that introspects any Gradio chat space (such as Hugging Face Spaces or custom deployments) and exposes a standards-compliant OpenAI /v1/chat/completions and /v1/models HTTP API.

Default demo space: https://tencent-hy3.hf.space

Features

  • Zero external dependencies: pure Go standard library (net/http, encoding/json, bufio, etc.).
  • Universal heuristic discovery engine:
    • Automatically interrogates /gradio_api/info, /config, and Hugging Face metadata without endpoint-specific hardcoding.
    • Detects Gradio runtime versions across v3, v4, v5, and v6.
    • Classifies application architecture into ChatInterface, Blocks (Chat), Blocks (Multimodal Chat), Interface, and Generic.
    • Disambiguation scoring engine evaluates candidate endpoints, filtering out UI resets, clears, retries, likes, and utility triggers to pinpoint primary conversational completion functions.
    • Correlates semantic parameter names from /gradio_api/info with component IDs from /config to reconstruct parameter mappings even when component labels are obfuscated.
  • Dual protocol support with auto-fallback:
    • Supports modern Gradio 4/5/6 /call SSE protocol with persistent session hashes.
    • Supports legacy Gradio 3 /run/predict and /api/predict direct execution protocol.
    • Instant zero-latency fallback from /call to /run/predict upon HTTP 404 or 405 status codes.
  • Tool calling support classification:
    • Automatically identifies tool calling mechanisms: native_slot, prompt_augmented_system, prompt_augmented_first_turn, or prompt_augmented_single_prompt.
    • Upstream tool call recovery intercepts tool_use_failed errors and extracts function names and JSON arguments.
    • Real-time sliding-window <tool_call> tag interceptor emits structured OpenAI tool call chunks.
  • Native Tencent Hunyuan 3 (tencent-hy3) support:
    • Full native zero-degradation handling for official spaces like https://tencent-hy3.hf.space.
    • Maps functions_json_str natively without polluting the system prompt.
    • Preserves multi-turn reasoning content and tool call history in standard OpenAI message schemas.
    • Maps reasoning_effort (no_think, low, high) directly to think_level.
  • Universal multi-turn handling:
    • Automatically formats conversation history into structured inputs when the space supports them.
    • Transparently composes multi-turn dialogue (System, User, Assistant) into single prompt inputs when the space only accepts a single message textbox.
    • Automatically pads hidden/State inputs (e.g. Gradio State components) to prevent backend argument count mismatches.
    • Automatically handles multimodal textbox inputs (MultimodalData with {text, files}) and space component defaults (radios, sliders, checkboxes).
  • Real-time streaming & accumulation filter:
    • Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames [[content, reasoning, tool_calls, history]]).
    • Emits standards-compliant chat.completion.chunk SSE events in real time.
  • Thinking & reasoning token separation:
    • Streams native reasoning chunks as delta.reasoning_content in real time.
    • Detects <think>...</think> tags in real time as fallback for standard spaces.
    • Separates reasoning into delta.reasoning_content (streaming) and message.reasoning_content (non-streaming).
    • Keeps content clean without tag leakage.
  • Built-in SOCKS5 proxy client:
    • Full RFC 1928 / RFC 1929 implementation with domain resolution (socks5h://), IPv4, IPv6, and username/password auth.
  • Dynamic space override:
    • Switch the target Gradio space on-the-fly per request using the X-Gradio-Space or X-Space-URL HTTP headers.
  • Fibonacci retry engine:
    • Resilient backoff retry mechanism (1s, 1s, 2s, 3s, 5s) for transient network hiccups with immediate break on 404/405 errors.

Heuristic discovery engine

The gateway implements an autonomous heuristic engine that discovers and configures the optimal completion path for any target Gradio space at startup.

Space introspection and version detection

Upon initialization, gr2gw inspects the space metadata:

  1. Queries /gradio_api/info and /config endpoints.
  2. Extracts the Gradio runtime version (v3, v4, v5, or v6).
  3. Detects API routing prefixes (e.g. /gradio_api on modern versions, or root on Gradio 3).

UI flavor classification

The engine classifies the application structure into architectural flavors:

  • ChatInterface: Standard Gradio chat interfaces equipped with chatbot, textbox, and optional additional inputs.
  • Blocks (Chat): Custom gr.Blocks layouts containing conversational components.
  • Blocks (Multimodal Chat): Blocks architectures featuring MultimodalTextbox components that accept {text, files} JSON payloads.
  • Interface: Classic input-output gr.Interface instances.
  • Generic: Spaces with custom or unclassified component topologies.

Candidate endpoint scoring

Gradio spaces frequently expose dozens of internal endpoints for UI actions (e.g. clearing text, retrying responses, voting/liking, adjusting sliders). The scoring algorithm identifies the true conversational endpoint by:

  • Penalizing non-conversational triggers (e.g. -600 for clear/reset/undo/retry/like endpoints).
  • Rewarding chat semantics (+150 for /chat, /predict, /generate, /respond).
  • Rewarding message inputs (+120 for Textbox or MultimodalTextbox).
  • Rewarding chat history slots (+80 for Chatbot or State components).
  • Rewarding generator and streaming dependencies (+50).

Dual protocol execution and auto-fallback

  • call protocol: Modern Gradio 4/5/6 execution via POST /call/{endpoint} returning an event ID, followed by GET /call/{endpoint}/{event_id} SSE streaming.
  • predict protocol: Gradio 3 and legacy execution via direct POST /run/predict or POST /api/predict.
  • Runtime failover: If a space returns HTTP 404 or 405 when calling the modern protocol, the gateway breaks immediately from the retry loop and falls back to /run/predict.

Tool calling support modes

The gateway evaluates available input components to determine how tool schemas and function calls should be delivered:

  • native_slot: The space provides a dedicated parameter slot for tool definitions (e.g. functions_json_str on Hunyuan 3). Function definitions are passed cleanly without prompt alteration.
  • prompt_augmented_system: The space provides a separate system_prompt input slot. Tool definitions and invocation schemas are injected directly into the system prompt.
  • prompt_augmented_first_turn: The space accepts chat history pairs but lacks a dedicated system prompt slot. Tool definitions are prepended to the user prompt on the first dialogue turn.
  • prompt_augmented_single_prompt: The space accepts only a single textbox input. Full multi-turn dialogue, tool definitions, and system guidance are synthesized into a single cohesive prompt.

Startup resolution diagnostics

Whenever the gateway starts or inspects a new space, it prints the complete resolution picture:

================================================================================
Gradio Space Resolution Picture
--------------------------------------------------------------------------------
Space URL:         https://tencent-hy3.hf.space
Title:             Hunyuan 3 Chat
Gradio Version:    5.29.0
UI Flavor:         Blocks (Chat)
Protocol:          call
API Prefix:        /gradio_api
Resolved Endpoint: /chat_fn
Function Index:    1
Primary Model:     hy3
Exposed Models:    hy3, hunyuan3, tencent/Hy3
History Format:    tuples
Tool Call Support: native_slot
Total Input Slots: 9
Input Slot Mappings:
  [0] Component ID 1    textbox (label="message")      -> message
  [1] Component ID 2    textbox (label="system")       -> system_prompt
  [2] Component ID 3    chatbot (label="chatbot")      -> history
  [3] Component ID 4    radio (label="think_level")    -> think_level
  [4] Component ID 5    slider (label="temperature")   -> temperature
  [5] Component ID 6    slider (label="max_tokens")    -> max_tokens
  [6] Component ID 7    slider (label="top_p")         -> top_p
  [7] Component ID 8    state (label="preserved")      -> preserved_thinking
  [8] Component ID 9    textbox (label="functions")    -> functions_json_str
================================================================================

Gateway status endpoint

Making a GET request to / returns a JSON summary of the running gateway and the discovered space profile:

curl http://localhost:8080/

Response:

{
  "name": "gr2gw",
  "status": "ready",
  "space_url": "https://tencent-hy3.hf.space",
  "gradio_version": "5.29.0",
  "flavor": "Blocks (Chat)",
  "protocol": "call",
  "tool_call_mode": "native_slot",
  "models": ["hy3", "hunyuan3", "tencent/Hy3"]
}

Installation

go install code.luxferre.top/luxferre/gr2gw@latest

Build

make build

Binary will be compiled to bin/gr2gw.

To run tests:

make test

Usage

Quick start

Run with the default space (https://tencent-hy3.hf.space):

./bin/gr2gw -port 8080

Target any other Gradio space:

./bin/gr2gw -space https://lucasmarchettidelima-digital-twin.hf.space -port 8080

With SOCKS5 proxy:

./bin/gr2gw -space https://tencent-hy3.hf.space -socks socks5://127.0.0.1:1080

CLI flags

Flag Default Description
-space, -url, -endpoint https://tencent-hy3.hf.space Target Gradio space URL
-model "" Exposed model name override (default: auto-detected)
-port 8080 Port to listen on
-host 0.0.0.0 Host interface to bind to
-socks, -proxy, -socks5 "" SOCKS5 proxy URL (socks5://user:pass@host:port)
-user-agent, -ua Firefox string Custom User-Agent header
-timeout 300 Upstream request timeout in seconds

Environment variables

  • GRADIO_SPACE_URL: default Gradio space URL fallback.
  • ALL_PROXY, SOCKS5_PROXY, SOCKS_PROXY: default SOCKS5 proxy URL fallback.

API examples

List models

curl http://localhost:8080/v1/models

Response:

{
  "object": "list",
  "data": [
    {
      "id": "hy3",
      "object": "model",
      "created": 1788756307,
      "owned_by": "gradio"
    },
    {
      "id": "hunyuan3",
      "object": "model",
      "created": 1788756307,
      "owned_by": "gradio"
    },
    {
      "id": "tencent/Hy3",
      "object": "model",
      "created": 1788756307,
      "owned_by": "gradio"
    }
  ]
}

Chat completions (non-streaming)

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hy3",
    "messages": [
      {"role": "user", "content": "What is the capital of France?"}
    ]
  }'

Response:

{
  "id": "chatcmpl-16425f9d-c350-47a1-9a6d-e9ce10871545",
  "object": "chat.completion",
  "created": 1788756310,
  "model": "hy3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of France is Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}

Chat completions (streaming)

curl -N http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hy3",
    "messages": [
      {"role": "user", "content": "Count from 1 to 5."}
    ],
    "stream": true
  }'

Tool calling (turn 1)

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hy3",
    "messages": [
      {"role": "user", "content": "What is the weather in Tokyo?"}
    ],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather in a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string", "description": "City name"}
          },
          "required": ["location"]
        }
      }
    }]
  }'

Response:

{
  "id": "chatcmpl-32727c62-ef2a-4866-855b-f1c7ec2b8023",
  "object": "chat.completion",
  "created": 1788756322,
  "model": "hy3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_3d4c016a",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\":\"Tokyo\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}

Tool response submission (turn 2)

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hy3",
    "messages": [
      {"role": "user", "content": "What is the weather in Tokyo?"},
      {
        "role": "assistant",
        "tool_calls": [
          {
            "id": "call_3d4c016a",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\":\"Tokyo\"}"
            }
          }
        ]
      },
      {
        "role": "tool",
        "tool_call_id": "call_3d4c016a",
        "content": "{\"temperature\": 20, \"condition\": \"sunny\"}"
      }
    ],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather in a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string", "description": "City name"}
          },
          "required": ["location"]
        }
      }
    }]
  }'

Response:

{
  "id": "chatcmpl-4903ba12-f12b-4cd3-a801-7290bc91a421",
  "object": "chat.completion",
  "created": 1788756335,
  "model": "hy3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The current weather in Tokyo is sunny with a temperature of 20 °C."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}

Dynamic target space override

Override the target space per request without restarting the server:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "X-Gradio-Space: https://ericsqin-hy3.hf.space" \
  -d '{
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Credits

Created by Luxferre in 2026, released into the public domain with no warranties.