Gateway allowlist request: Fusion AI Gateway — production wholesale broker (DeepSeek Flash / MiniMax / Kimi) #1637
Operator
- Name: Aung Myat Moe
- GitHub: @theaungmyatmoe
- Project: Fusion AI Gateway (
https://api.fusioncode.app) — production OpenAI-compatible wholesale gateway - Contact: via this GitHub issue
Requested creator address
Public key (compressed secp256k1, hex):
Please consider adding this address to devshard_escrow_params.allowed_creator_addresses.
The key is dedicated exclusively to devshard escrow creation. Per the self-hosted gateway guide, the creator will be funded only after allowlist membership is confirmed.
Models
deepseek-ai/DeepSeek-V4-Flash-0731— primary (cache-heavy agent workload)MiniMaxAI/MiniMax-M2.7— secondarymoonshotai/Kimi-K2.6— secondary / evaluation
Use case
We operate a production wholesale gateway already serving real OpenAI-compatible traffic at api.fusioncode.app. We broker DeepSeek V4 Flash, MiniMax M2.7, and Kimi K2.6 to downstream clients (agent infrastructure, high-volume background processing).
We currently route through OpenBroker, but need the native devshard creator path because:
- we require direct GNK settlement and self-custody for a committed wholesale volume (multi-billion tokens/day pipeline);
- we need control over escrow pooling, rotation, settlement, and capacity-aware routing for a sustained 60–100 RPS baseline with burst support;
- we need the network per-token price rather than a broker layer between us and the chain;
- broker-side concurrency limits have surfaced under sustained load.
This request is for a production self-hosted gateway, not a broker-directory listing.
Initial deployment plan
- Gateway-only deployment (
libermans/gonka-devshard-proxy) on an existing Linux VPS - Public node endpoints (no full chain node),
GATEWAY_MAX_CONCURRENT_REQUESTS=512 - Multiple devshards pooled for throughput and epoch rotation; capacity-aware limits enabled
- Reverse proxy in front; creator key dedicated and not reused
- Staged concurrency ramp, settlement/refund verification, then production cutover
Validation and contribution plan
- One manually managed escrow + deterministic functional checks.
- Concurrency ramp: 1 / 3 / 10 / 25 / 50 / 100 / 200.
- Measure success rate, HTTP error distribution, TTFT, p50/p95/p99 latency, output throughput, settlement/refund behavior.
- Enable multi-escrow pooling + rotation only after single-escrow verification.
- Share anonymized capacity and reliability results with the Gonka community.
Governance and operations commitments
- We understand allowlist inclusion is an on-chain governance decision and is not guaranteed.
- We will respond to maintainer and host questions promptly.
- We will not fund or open escrows until the address is confirmed on the allowlist.
- We will start with staged private traffic, not an immediate public endpoint.
- We will publish operational findings and adjust limits if governance or operators request it.
Related
- Open discussion of gateway economics and cache-served token pricing: https://github.com/gonka-ai/gonka/discussions/1636
- Telemetry PR (cache metadata passthrough): https://github.com/gonka-ai/gonka/pull/1633
- Precedent request: #1479 (Knyazev AI)
🔄 Auto-synced from Issue #1637 every hour.