feat: implement universal Gradio to OpenAI proxy gateway
This commit is contained in:
@@ -0,0 +1,251 @@
|
||||
# gr2gw: Universal Gradio to OpenAI LLM Gateway
|
||||
|
||||
A zero-dependency, high-performance Go proxy server that introspects any Gradio chat space (such as Hugging Face Spaces or custom deployments) and exposes a standards-compliant OpenAI `/v1/chat/completions` and `/v1/models` HTTP API.
|
||||
|
||||
Default demo space: `https://ghost2513-openai-gpt-oss-120b.hf.space`
|
||||
|
||||
---
|
||||
|
||||
## Features
|
||||
|
||||
- **Zero External Dependencies**: Pure Go standard library (`net/http`, `encoding/json`, `bufio`, etc.).
|
||||
- **Automatic Space Introspection**: Dynamically queries `/gradio_api/info`, `/config`, and Hugging Face space metadata to discover models, endpoints, and input parameter mappings.
|
||||
- **Universal Multi-turn Handling**:
|
||||
- Automatically formats conversation history into structured inputs when the space supports them.
|
||||
- Transparently composes multi-turn dialogue (`System`, `User`, `Assistant`) into single prompt inputs when the space only accepts a single message textbox.
|
||||
- Automatically pads hidden/State inputs (e.g. Gradio State components) to prevent backend argument count mismatches.
|
||||
- **Real-Time Streaming & Accumulation Filter**:
|
||||
- Automatically computes token deltas from cumulative or incremental Gradio SSE output streams.
|
||||
- Emits standards-compliant `chat.completion.chunk` SSE events in real time.
|
||||
- **Thinking & Reasoning Token Separation**:
|
||||
- Detects `<think>...</think>` tags in real time.
|
||||
- Separates reasoning into `delta.reasoning_content` (streaming) and `message.reasoning_content` (non-streaming).
|
||||
- Keeps `content` clean without tag leakage.
|
||||
- **Full Tool Calling & Function Interception**:
|
||||
- Formats schemas into system prompts with strict function calling instructions.
|
||||
- **`StreamToolCallFilter`**: Stateful sliding-window filter that prevents `<tool_call>` tags from leaking into `delta.content`. Emits structured OpenAI `delta.tool_calls` chunks and sets `finish_reason: "tool_calls"`.
|
||||
- Seamlessly maintains multi-turn context when tool results are submitted back via `role: "tool"`.
|
||||
- **Built-in SOCKS5 Proxy Client**:
|
||||
- Full RFC 1928 / RFC 1929 implementation with domain resolution (`socks5h://`), IPv4, IPv6, and username/password auth.
|
||||
- **Dynamic Space Override**:
|
||||
- Switch the target Gradio space on-the-fly per request using the `X-Gradio-Space` or `X-Space-URL` HTTP headers.
|
||||
- **Fibonacci Retry Engine**:
|
||||
- Resilient backoff retry mechanism (1s, 1s, 2s, 3s, 5s) for transient network hiccups.
|
||||
|
||||
---
|
||||
|
||||
## Build
|
||||
|
||||
```bash
|
||||
make build
|
||||
```
|
||||
|
||||
Binary will be compiled to `bin/gr2gw`.
|
||||
|
||||
To run tests:
|
||||
```bash
|
||||
make test
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Usage
|
||||
|
||||
### Quick Start
|
||||
|
||||
Run with the default space (`https://ghost2513-openai-gpt-oss-120b.hf.space`):
|
||||
```bash
|
||||
./bin/gr2gw -port 8080
|
||||
```
|
||||
|
||||
Target any other Gradio space:
|
||||
```bash
|
||||
./bin/gr2gw -space https://ericsqin-hy3.hf.space -port 8080
|
||||
```
|
||||
|
||||
With SOCKS5 proxy:
|
||||
```bash
|
||||
./bin/gr2gw -space https://ghost2513-openai-gpt-oss-120b.hf.space -socks socks5://127.0.0.1:1080
|
||||
```
|
||||
|
||||
### CLI Flags
|
||||
|
||||
| Flag | Default | Description |
|
||||
|------|---------|-------------|
|
||||
| `-space`, `-url` | `https://ghost2513-openai-gpt-oss-120b.hf.space` | Target Gradio space URL |
|
||||
| `-port` | `8080` | Port to listen on |
|
||||
| `-host` | `0.0.0.0` | Host interface to bind to |
|
||||
| `-socks`, `-proxy`, `-socks5` | `""` | SOCKS5 proxy URL (`socks5://user:pass@host:port`) |
|
||||
| `-user-agent`, `-ua` | Firefox string | Custom User-Agent header |
|
||||
| `-timeout` | `300` | Upstream request timeout in seconds |
|
||||
|
||||
### Environment Variables
|
||||
|
||||
- `GRADIO_SPACE_URL`: Default Gradio space URL fallback.
|
||||
- `ALL_PROXY`, `SOCKS5_PROXY`, `SOCKS_PROXY`: Default SOCKS5 proxy URL fallback.
|
||||
|
||||
---
|
||||
|
||||
## API Examples
|
||||
|
||||
### List Models
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/models
|
||||
```
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "openai/gpt-oss-120b",
|
||||
"object": "model",
|
||||
"created": 1788756307,
|
||||
"owned_by": "gradio"
|
||||
},
|
||||
{
|
||||
"id": "gpt-oss-120b",
|
||||
"object": "model",
|
||||
"created": 1788756307,
|
||||
"owned_by": "gradio"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Chat Completions (Non-Streaming)
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "openai/gpt-oss-120b",
|
||||
"messages": [
|
||||
{"role": "user", "content": "What is the capital of France?"}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-16425f9d-c350-47a1-9a6d-e9ce10871545",
|
||||
"object": "chat.completion",
|
||||
"created": 1788756310,
|
||||
"model": "openai/gpt-oss-120b",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "Paris is the capital of France."
|
||||
},
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 0,
|
||||
"completion_tokens": 0,
|
||||
"total_tokens": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Chat Completions (Streaming)
|
||||
|
||||
```bash
|
||||
curl -N http://localhost:8080/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "openai/gpt-oss-120b",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Count from 1 to 5."}
|
||||
],
|
||||
"stream": true
|
||||
}'
|
||||
```
|
||||
|
||||
### Tool Calling
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "openai/gpt-oss-120b",
|
||||
"messages": [
|
||||
{"role": "user", "content": "What is the weather in Tokyo?"}
|
||||
],
|
||||
"tools": [{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get current weather in a location",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"location": {"type": "string", "description": "City name"}
|
||||
},
|
||||
"required": ["location"]
|
||||
}
|
||||
}
|
||||
}]
|
||||
}'
|
||||
```
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-32727c62-ef2a-4866-855b-f1c7ec2b8023",
|
||||
"object": "chat.completion",
|
||||
"created": 1788756322,
|
||||
"model": "openai/gpt-oss-120b",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": null,
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_3d4c016a",
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"arguments": "{\"location\":\"Tokyo\"}"
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
"finish_reason": "tool_calls"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 0,
|
||||
"completion_tokens": 0,
|
||||
"total_tokens": 0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Dynamic Target Space Override
|
||||
|
||||
Override the target space per request without restarting the server:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "X-Gradio-Space: https://ericsqin-hy3.hf.space" \
|
||||
-d '{
|
||||
"messages": [
|
||||
{"role": "user", "content": "Hello!"}
|
||||
]
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## License
|
||||
|
||||
Released into the public domain under Creative Commons Zero (CC0) or Unlicense.
|
||||
Reference in New Issue
Block a user