gr2gw: Universal Gradio to OpenAI LLM gateway

A zero-dependency, high-performance Go proxy server that introspects any Gradio chat space (such as Hugging Face Spaces or custom deployments) and exposes a standards-compliant OpenAI /v1/chat/completions and /v1/models HTTP API.

Default demo space: https://ghost2513-openai-gpt-oss-120b.hf.space

Features

  • Zero external dependencies: Pure Go standard library (net/http, encoding/json, bufio, etc.).
  • Automatic space introspection: Dynamically queries /gradio_api/info, /config, and Hugging Face space metadata to discover models, endpoints, and input parameter mappings.
  • Native Tencent Hunyuan 3 (tencent-hy3) support:
    • Full native zero-degradation handling for official spaces like https://tencent-hy3.hf.space.
    • Maps functions_json_str natively without polluting the system prompt.
    • Preserves multi-turn reasoning content and tool call history in standard OpenAI message schemas.
    • Maps reasoning_effort (no_think, low, high) directly to think_level.
  • Universal multi-turn handling:
    • Automatically formats conversation history into structured inputs when the space supports them.
    • Transparently composes multi-turn dialogue (System, User, Assistant) into single prompt inputs when the space only accepts a single message textbox.
    • Automatically pads hidden/State inputs (e.g. Gradio State components) to prevent backend argument count mismatches.
  • Real-time streaming & accumulation filter:
    • Automatically computes token deltas from cumulative or incremental Gradio SSE output streams (including 2D Hy3 frames [[content, reasoning, tool_calls, history]]).
    • Emits standards-compliant chat.completion.chunk SSE events in real time.
  • Thinking & reasoning token separation:
    • Streams native reasoning chunks as delta.reasoning_content in real time.
    • Detects <think>...</think> tags in real time as fallback for standard spaces.
    • Separates reasoning into delta.reasoning_content (streaming) and message.reasoning_content (non-streaming).
    • Keeps content clean without tag leakage.
  • Full tool calling & function interception:
    • Formats schemas into native functions_json_str (Hy3) or system prompts (standard spaces).
    • StreamToolCallFilter: Stateful sliding-window filter that prevents <tool_call> tags from leaking into delta.content. Emits structured OpenAI delta.tool_calls chunks and sets finish_reason: "tool_calls".
    • Seamlessly maintains multi-turn context when tool results are submitted back via role: "tool".
  • Built-in SOCKS5 proxy client:
    • Full RFC 1928 / RFC 1929 implementation with domain resolution (socks5h://), IPv4, IPv6, and username/password auth.
  • Dynamic space override:
    • Switch the target Gradio space on-the-fly per request using the X-Gradio-Space or X-Space-URL HTTP headers.
  • Fibonacci retry engine:
    • Resilient backoff retry mechanism (1s, 1s, 2s, 3s, 5s) for transient network hiccups.

Installation

go install code.luxferre.top/luxferre/gr2gw@latest

Build

make build

Binary will be compiled to bin/gr2gw.

To run tests:

make test

Usage

Quick start

Run with the default space (https://ghost2513-openai-gpt-oss-120b.hf.space):

./bin/gr2gw -port 8080

Target any other Gradio space:

./bin/gr2gw -space https://ericsqin-hy3.hf.space -port 8080

With SOCKS5 proxy:

./bin/gr2gw -space https://ghost2513-openai-gpt-oss-120b.hf.space -socks socks5://127.0.0.1:1080

CLI flags

Flag Default Description
-space, -url https://ghost2513-openai-gpt-oss-120b.hf.space Target Gradio space URL
-port 8080 Port to listen on
-host 0.0.0.0 Host interface to bind to
-socks, -proxy, -socks5 "" SOCKS5 proxy URL (socks5://user:pass@host:port)
-user-agent, -ua Firefox string Custom User-Agent header
-timeout 300 Upstream request timeout in seconds

Environment variables

  • GRADIO_SPACE_URL: default Gradio space URL fallback.
  • ALL_PROXY, SOCKS5_PROXY, SOCKS_PROXY: default SOCKS5 proxy URL fallback.

API examples

List models

curl http://localhost:8080/v1/models

Response:

{
  "object": "list",
  "data": [
    {
      "id": "openai/gpt-oss-120b",
      "object": "model",
      "created": 1788756307,
      "owned_by": "gradio"
    },
    {
      "id": "gpt-oss-120b",
      "object": "model",
      "created": 1788756307,
      "owned_by": "gradio"
    }
  ]
}

Chat completions (non-streaming)

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "What is the capital of France?"}
    ]
  }'

Response:

{
  "id": "chatcmpl-16425f9d-c350-47a1-9a6d-e9ce10871545",
  "object": "chat.completion",
  "created": 1788756310,
  "model": "openai/gpt-oss-120b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Paris is the capital of France."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}

Chat completions (streaming)

curl -N http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "Count from 1 to 5."}
    ],
    "stream": true
  }'

Tool calling

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "What is the weather in Tokyo?"}
    ],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather in a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string", "description": "City name"}
          },
          "required": ["location"]
        }
      }
    }]
  }'

Response:

{
  "id": "chatcmpl-32727c62-ef2a-4866-855b-f1c7ec2b8023",
  "object": "chat.completion",
  "created": 1788756322,
  "model": "openai/gpt-oss-120b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_3d4c016a",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\":\"Tokyo\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}

Dynamic target space override

Override the target space per request without restarting the server:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "X-Gradio-Space: https://ericsqin-hy3.hf.space" \
  -d '{
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Credits

Created by Luxferre in 2026, released into the public domain with no warranties.

S
Description
Turn any suitable keyless Gradio space into an LLM gateway
Readme
1.2 MiB
Languages
Go 99.9%