FAQ: Multi-Model Access, Smart Routing, Cost Governance & Audit

An enterprise LLM gateway is a layer that puts multiple model providers behind one OpenAI-compatible endpoint, with routing, budgets, permissions, and audit applied centrally. This page answers, by theme, what enterprises actually ask before adopting one: getting started and integration, routing and reliability, governance and cost, compliance and data boundaries, and where the line runs against self-hosted or resale alternatives. Each answer is short enough to quote on its own.

How do I get started and integrate?

How do I call multiple model providers through one OpenAI-compatible API?

Point your application base_url at the smaapi gateway (SMA) and keep using the OpenAI SDK and request format. The gateway translates protocols and handles provider authentication; you pick the target model via the model parameter or let routing policy decide. No per-provider integration code.

How do I get started with SMA and open an API account?

Register in the console, create an API key on the key-management page, pick a model group, and enable it. Top up via online Alipay (instant) or corporate bank transfer (with a technical-service invoice). Keep your OpenAI SDK and point base_url at the gateway to start calling. Experienced teams self-serve in minutes; if you prefer help, 24/7 technical support via WeCom is on standby.

Do I need to be technical to integrate? Is there help?

Both work. Experienced developers self-integrate — SMA is OpenAI-compatible, so you change base_url and key. If you would rather have help, 24/7 technical support via WeCom assists with installation and integration until you are connected.

How do I pay, and can you provide an invoice (fapiao)?

Top up via online Alipay (instant) or corporate bank transfer; we settle in both CNY and USD (USD is available for overseas entities — exchange rate and receiving account are specified in the contract). Billing is usage-based, with real-time usage and cost in the console. Corporate payments come with a technical-service invoice (fapiao) — the invoicing entity, invoice type, and VAT rate are confirmed with the contract and documentation during procurement; we only commit to terms we can put in writing.

Which models does SMA support?

SMA uses an OpenAI-compatible protocol as its unified access layer and connects to mainstream international and Chinese model providers. The authoritative list is the live, health-checked roster published in the console — only actually connected models are listed; planned ones are labeled separately.

How do routing, governance, and reliability work?

What logic does smart routing use to dispatch requests?

Routing combines three signal groups: task characteristics (business tags and complexity), model state (availability, latency, cost), and enterprise policy (budgets, compliance boundaries, manual calibration rules). Scoring proposes candidates; policy and safety boundaries constrain the final choice, with automatic failover.

How does cost governance and usage auditing work?

Budgets, quotas, and rate limits are set per team, application, or key at the gateway layer, and every call is metered through the same pipeline. Each API key is returned only once at creation and stored as a hash plus mask; exceeding quota or rate limit is blocked (429, with no billing and no model call). Audit logs record caller, model, usage, cost, and time, ready to aggregate and reconcile per key, project, or model. Because all traffic passes through the gateway, the numbers are complete rather than a patchwork of provider dashboards.

What happens when a model endpoint fails?

Failed requests are handled by error type: recoverable errors (rate limit, timeout, 5xx, empty response) switch to a candidate model and preserve context via the state header; authentication or content-policy errors do not switch, to keep incidents from spreading. Fallback order is configurable per business priority, transparent to applications, and fully audited. Active real-time health monitoring and preemptive switching are provided by SMA Edge (the physical gateway).

How do you keep access to Claude, GPT, and other overseas models stable?

Stability comes from redundancy, governance, and monitoring, not a single point. In software: multiple models back each other up, and when a model endpoint returns rate-limit, timeout, or failure, SMA switches to a candidate and preserves context via the state header — transparent to the application, fully audited, reducing single-point impact. In hardware: SMA Edge (the physical gateway) runs active real-time health monitoring of model endpoints with preemptive switching — this requires the SMA gateway hardware. The application always sees one unchanging OpenAI-compatible egress, so provider or route changes do not affect your integration. Specific availability and SLA follow your enterprise contract; we make no absolute "never down" guarantee.

How are compliance and data boundaries defined?

How can a company in China use the Claude API through proper channels?

Two layers: channel and scenario. On channel, access should run through legitimate enterprise-grade routes with a full contract chain and audit records — not personal accounts or resale of unknown origin. On scenario, internal business use and public-facing services are different regimes: services offered to the Chinese public must use government-registered models. SMA provides unified access, routing, and audit on top of legitimate channels, and can route public-facing traffic to registered domestic models.

What are your upstream channels?

Mainstream overseas closed-source models (such as GPT) are accessed through the commercial platforms of cloud providers — GPT runs on the Microsoft Azure commercial service, with the contracting entity and accountability written into contracts, technical-service invoicing, and full call auditing; other mainstream models are being integrated to the same standard. Legitimacy is verifiable: contracting entity, invoicing, auditability, and data boundary. Available models and channels follow the live data and the contract.

Does our data leave China? How is it sanitized?

When calling international models, prompts are transmitted to endpoints outside China — that is a fact we do not paper over. The gateway provides sensitive-data detection and sanitization at the egress, and every cross-border call enters the audit log. For workloads involving personal information or important data, routing policy should follow your own cross-border assessment: SMA can route such requests to registered domestic models so that data stays in-country.

Can public-facing products use international models?

The honest answer: under current regulation, generative AI services offered to the public in China must use registered models. Public-facing traffic should therefore be served by registered domestic models — SMA routes it there automatically — while international models fit internal business scenarios. One gateway, two scopes, each within its boundary.

How does it compare to the alternatives?

What is an enterprise LLM gateway?

A unified access and governance layer between your applications and multiple LLM providers: one OpenAI-compatible API on top, smart routing, cost and permission governance, and full audit trails underneath. It turns multi-model usage from scattered integrations into one controlled egress.

How is it different from a self-hosted relay or proxy tool?

Relay tools (OneAPI-style projects) answer "can I reach many models" and fit individuals and small teams. An enterprise gateway adds organization-level governance: fine-grained permissions and key custody, budgets and usage audit, full-chain call logs, failover, and a clear compliance boundary. One is a personal tool; the other is infrastructure.

How do I choose the best enterprise LLM gateway?

Evaluate against six criteria: OpenAI protocol compatibility, routing configurability and failover behavior, governance granularity (team / app / key), audit completeness, honesty of model coverage (verifiable connectivity data, not claims), and the data boundary including a self-hosted option. Weigh them by your compliance posture and traffic profile.

How is this different from an API relay, in terms of channel and accountability?

Relay-style services often sit on personal account pools or resale of unknown origin, with no contracting entity to hold accountable when things break. An enterprise gateway runs on legitimate channels with the contracting party, SLA, audit, and responsibility written down. What an enterprise actually procures is not an endpoint — it is someone accountable and books that reconcile.

Which of these answers rest on primary sources?

The compliance answers above rest on public texts you can check, not on our reading of them. Take the question asked most often — whether a security assessment and algorithm filing are required. Article 17 of China's Interim Measures for the Management of Generative AI Services reads, verbatim:

"提供具有舆论属性或者社会动员能力的生成式人工智能服务的,应当按照国家有关规定开展安全评估,并按照《互联网信息服务算法推荐管理规定》履行算法备案和变更、注销备案手续。"

— Interim Measures for the Management of Generative AI Services, Article 17. Working translation: providers of generative AI services with public-opinion attributes or social-mobilisation capacity must carry out a security assessment under the relevant state rules and complete algorithm filing (and any change or cancellation of that filing) under the Provisions on the Management of Algorithmic Recommendations in Internet Information Services. Source: Cyberspace Administration of China (the Chinese text is authoritative).

The obligation attaches to the shape of the service, not to which vendor's model is behind it: whether it carries public-opinion attributes or social-mobilisation capacity, and whether it is offered to the public within mainland China. What a gateway can do is route public-facing traffic to registered domestic models and keep calls and data boundaries logged for review. Whether a specific workload falls under this article is a determination each enterprise makes with its own counsel — we do not make that call on your behalf. The companion provision on scope (Article 2(3)) is quoted on the transparency page.

References

About the name: smaapi (the SMA gateway) is the enterprise AI gateway built by Slime Mould Tech — SMA stands for Slime Mould Architecture. smaapi is unrelated to the simple moving average indicator in finance, to the solar inverter vendor SMA Solar Technology AG, or to the SMA coaxial connector standard that share the acronym.