Search

Secure Deployment: VPC, On-Premises Deployment, and Isolated Environments

Enterprise-Grade AI Gateway

Built-in enterprise-grade governance and observability capabilities

Securely manage and govern more than 1,600 AI models through a unified AI Gateway. With policy controls, real-time monitoring, and intelligent optimization, help enterprises reduce AI usage costs by up to 30%.

Designed for large-scale AI applications in real-world production environments

Processed monthly
Over 10 billion requests

Provide high-throughput, highly scalable inference services for production-grade AI applications.

Average decrease
30% AI Cost

Intelligent routing, batch processing, and budget control mechanisms significantly reduce token waste and optimize inference costs.

99.99%
Service Availability

Centralized failover, intelligent routing, and safety guardrail mechanisms ensure that AI applications continue to operate stably even if a model service provider experiences a failure.

1600+
Model Support

The Unified AI Gateway connects and manages various mainstream AI models, enabling centralized access and unified governance.

AI Gateway: A Unified Integration Layer for Large Model APIs

Simplify your generative AI tech stack by integrating all major models through a single AI gateway.

  • Through a single AI Gateway API, you can connect to leading models such as OpenAI, Claude, Gemini, and DeepSeek, as well as over 250 large language models (LLMs).
  • Supports various model types, including chat, text generation (completion), vector embedding, and reranking.
  • Centrally manage API keys and team authentication to enable unified access control.
  • Easily orchestrate multi-model workloads within your infrastructure to enable flexible invocation and collaboration.

AI Gateway Observability

Gain real-time insight into the operational status, usage costs, and compliance status of the AI gateway to achieve end-to-end visibility and management.

  • Monitor key metrics in the real-time monitoring system, such as token usage, response latency, error rates, and request volume.
  • Centrally store and view complete request and response logs to meet audit and compliance requirements and streamline the troubleshooting process.
  • Add metadata tags such as user ID, team, and environment to traffic to gain more granular operational insights.
  • Quickly filter logs and metrics by model, team, or region to pinpoint the root cause of issues and accelerate troubleshooting.
AI 網關可觀測性儀表板

AI Gateway Quotas and Access Control

Achieve AI governance, cost control, and risk protection through enterprise-level strategy management capabilities.

  • Set rate limits for users, services, or interfaces.
  • Configure quotas based on cost or token consumption using metadata rules.
  • Use Role-Based Access Control (RBAC) to implement resource isolation and permission management.
  • Achieve large-scale governance by centrally managing service accounts and agent workloads through unified rules.
AI 網關配額與存取控制介面

Low-Latency Inference

Leveraging high-performance AI gateway infrastructure to deliver exceptional inference performance for mission-critical scenarios.

  • Even under high-concurrency enterprise-level workloads, internal processing latency remains below 3 ms.
  • Flexibly scale computing power to easily handle traffic spikes and large-scale inference tasks.
  • Provides stable and predictable response times for instant messaging, RAG applications, and AI assistants.
  • Deploy nodes close to the inference layer to minimize network latency and reduce link loss.
Control Plane 架構圖

By deploying the AI gateway directly within the production inference pipeline, its low-latency architecture provides comprehensive governance and control capabilities without sacrificing performance.

AI Gateway Intelligent Routing and Failover

Through an intelligent traffic routing mechanism, business continuity is ensured even if a model fails.

  • Automatically select the fastest available language model based on response latency.
  • It uses a weighted load-balancing strategy to intelligently distribute traffic, thereby achieving greater reliability and scalability.
  • When a model call fails, the system automatically switches to a backup model to continue processing the request.
  • Supports region-based intelligent routing to meet data compliance and availability requirements in different regions.
Rate Limit 配置程式碼示意圖

This intelligent routing system can effectively prevent business disruptions caused by service interruptions or sudden spikes in latency in a single model,Ensure that enterprise AI applications remain online at all timesThe

Support for the self-hosted model

Comprehensive control over the deployment and execution environment for open-source models.

  • You can quickly integrate with mainstream open-source models such as LLaMA, Mistral, and Falcon without modifying the SDK.
  • Fully compatible with mainstream inference frameworks such as vLLM, SGLang, KServe, and Triton.
  • Helm-based automated scaling, GPU scheduling, and deployment management significantly simplify operations and maintenance.
  • Supports deploying and running models in VPCs, on-premises data centers, hybrid cloud environments, and physically isolated environments.
自託管模型部署介面示意圖

Integration of the AI Gateway and MCP

Build secure and reliable agent workflows with native MCP (Model Context Protocol) support.

  • Quickly integrate with enterprise-grade tools such as Slack, GitHub, Confluence, Datadog, and more.
  • Easily register and manage your internal MCP Server to reduce integration complexity.
  • Apply OAuth 2.0, RBAC, and metadata policy controls uniformly to every tool invocation.
MCP Servers 整合介面示意圖

AI Gateway Security Barrier

Build secure, trustworthy AI applications through configurable security safeguards and policy control mechanisms.

  • Seamlessly integrate and enforce enterprise-customized security policies, including features such as personal identifiable information (PII) filtering and malicious content detection.
  • Flexibly customize AI gateway security policy rules in accordance with corporate compliance requirements and security standards.
AI 網關安全護欄介面示意圖

Enterprise-Grade AI Gateway

Deploy a secure AI gateway designed specifically for enterprise environments to ensure that your data and models always run within your cloud environment or on-premises infrastructure.

  • Seamlessly integrate and enforce enterprise-customized security policies, including features such as personal identifiable information (PII) filtering and malicious content detection.
  • Flexibly customize AI gateway security policy rules in accordance with corporate compliance requirements and security standards.
HIPAA GDPR AICPA SOC 合規認證標章

Compliance and Security

Compliant with international standards such as SOC 2, HIPAA, and GDPR, it provides enterprises with comprehensive data security and privacy protection capabilities.

Governance and Access Control

Supports single sign-on (SSO), role-based access control (RBAC), and audit logs to enable unified identity and access management.

Enterprise-Level Support and Reliability

We provide 24/7 technical support and ensure the stable operation of mission-critical business systems through an SLA (Service Level Agreement).

Deploy in any environment

Supports deployment in VPCs, on-premises data centers, physically isolated environments, and multi-cloud architectures.

Your data remains under your control at all times and never leaves your organization’s boundaries. Regardless of the environment in which TrueFoundry is deployed, you benefit from full data sovereignty, environment isolation capabilities, and enterprise-grade compliance assurance.

多雲與本地部署架構示意圖

Frequently Asked Questions

Technical support and paid services to drive your project to the limit.

AI Gateway is a middleware platform specifically designed to connect, manage, and deploy artificial intelligence models and services. It acts as an intermediary between large language models (LLMs) and enterprise applications, establishing a secure, efficient, and unified channel for invoking these models. For example, models such as OpenAI GPT and Claude can be uniformly accessed and centrally managed through AI Gateway.

AI Gateway helps enterprises address challenges such as complex model integration, fragmented permissions, unmanageable costs, and a lack of governance capabilities, thereby enabling the large-scale deployment of AI applications.

AI Gateway acts as an intermediary between enterprise applications and model service providers, intelligently processing model requests. It offers core capabilities such as traffic routing, authentication, failover, and access control, ensuring that enterprise systems can access the necessary models and tools in a stable and efficient manner.

AI Gateway provides enterprises with a unified AI service management platform. Its core benefits include:

  • Unified Access to Multiple AI Models and Services
  • Provide identity authentication and access control mechanisms
  • Meet corporate security and compliance requirements
  • Implement usage monitoring and budget management
  • Supports intelligent load balancing and failover
  • Enforcing Data Usage and Security Policies
  • Supports horizontal scaling and rapid integration of new models

With AI Gateway, enterprises can manage their AI infrastructure in a more secure, efficient, and controllable manner.

AI Gateway provides unified access across models, intelligent routing, and automatic failover capabilities. Its core features include:

  • Governance and Security: Achieve enterprise-level governance and security through identity authentication, access control, and policy enforcement mechanisms.
  • Cost Optimization: Reduce AI usage costs through rate limiting, token quota management, and budget control mechanisms.
  • Comprehensive Observability: Track usage, performance metrics, and operating costs in real time.
  • Agent Workflow Support: Supports complex agent workflow orchestration and the execution of multi-step tasks.

As a unified control plane for enterprise AI infrastructure, AI Gateway helps organizations scale their AI deployments in a more secure and efficient manner.

Hongke AI Gateway provides comprehensive AI service deployment and management capabilities, designed specifically for enterprise-level scenarios. Its key features include:

  • RBAC, OAuth 2.0, and API Keys: Multi-Layer Security Authentication
  • Request Rate Limiting and Intelligent Load Balancing
  • Automatic Failover Mechanism
  • Built-in security and content governance capabilities
  • Comprehensive logging, analytics, and observability
  • Cloud Architecture Support
  • Real-time reasoning ability

For enterprises that need to balance security, governance, performance, and cost control, Hongke offers a flexible and scalable AI infrastructure solution.

Traditional API gateways are primarily responsible for forwarding and managing general network requests. AI gateways, on the other hand, are specifically designed for large language model (LLM) scenarios and can handle many tasks that traditional API gateways cannot perform efficiently, such as:

  • Token Statistics and Management
  • Prompt Cache
  • Model Routing
  • Model Failure Transition
  • AI Cost Control
  • Model Governance and Security Policy Enforcement

Therefore, the AI gateway is not only a traffic entry point but also a unified governance and control center for enterprise AI applications.

AI Gateway is deployed directly within the production inference chain between the application system and the model service.

As a unified control plane, it is responsible for:

  • Model Routing
  • Safety Controls
  • Permissions Management
  • Observability Monitoring
  • Cost Management
  • Agent and Tool Orchestration

The entire process provides unified AI management capabilities without requiring any changes to the application logic.

Support. The Enterprise AI Gateway supports not only hosted models but also self-hosted or open-source models such as LLaMA and Mistral.

These models can be deployed in a VPC, on-premises data center, hybrid cloud, or isolated environment, while sharing the same policy controls, security governance, and observability capabilities as hosted models. Enterprises can achieve unified operations and governance without having to build separate management systems for different models.

AI Gateway provides real-time usage monitoring, token-level tracking, quota management, and budget control capabilities. Additionally, AI Gateway continuously optimizes AI spending through the following methods:

  • Intelligent Model Routing
  • Request Cache
  • Automatic Failover
  • Traffic Scheduling Optimization

By reducing unnecessary calls to high-cost models, we effectively prevent inference costs from spiraling out of control and continuously improve resource utilization.

AI Gateway enables the unified implementation of enterprise data governance strategies, including:

  • Personal Identifiable Information (PII) Masking/De-identification
  • Request Filtering
  • Log Auditing and Access Control
  • Data Processing Rules Management

When deployed in a VPC, on-premises data center, or a standalone isolated environment, all sensitive data remains within the enterprise perimeter at all times and never leaves the organization’s control. This helps enterprises ensure data security while meeting industry regulatory and compliance requirements.

AI Gateway implements team-level isolation and governance through the following mechanisms:

  • Role-Based Access Control (RBAC)
  • Team-Level API Key Management
  • Quota Control
  • Usage Tracking and Auditing

Multiple teams can securely share model resources and infrastructure while maintaining independent permissions, cost transparency, and accountability.

Playground is an interactive development environment built on top of AI Gateway. Developers can quickly test different options before deploying models to production systems:

  • Large Language Models (LLMs)
  • Prompt
  • MCP Tools
  • Inference Configuration Parameters

In Playground, users can:

  • Select an integrated model from the Models page
  • Adjust parameters such as Temperature and Max Tokens
  • Configuring Streaming Output
  • Set up generation rules such as Stop Sequences

The system will display the following in real time:

  • Model Response Results
  • Token Usage
  • Response Delay Data

The development team can evaluate models and fine-tune parameters without writing code, significantly shortening the experimentation cycle.

Once the test results meet the requirements, the complete configuration—including the prompt, model, tools, security measures, and structured output definitions—can be saved as a reusable template and shared within the team.

Playground can also automatically generate code examples for frameworks such as the OpenAI SDK and LangChain, helping development teams quickly deploy validated solutions to production environments.

Through the Hongke AI Gateway, all model services and tools are unified into a single API layer. Development teams no longer need to manage SDKs, API endpoints, and access credentials for OpenAI, Anthropic, Amazon Bedrock, self-hosted models, and other model providers separately. Applications simply connect to a single Gateway Endpoint and use a single Gateway Key to access all model resources.

AI Gateway automatically handles model routing based on the configuration, so even if you replace a model or switch providers, you won't need to modify your application code.

This unified access capability extends beyond models to include:

  • MCP (Model Context Protocol) Tool Ecosystem
  • A2A (Agent-to-Agent) Protocol
  • Agent Workflow

This enables unified orchestration and management of models, tools, and agents.

For developers, this means:

  • A Simpler Approach to System Integration
  • A More Consistent Development Experience
  • A Clearer Safety Management Model

The "Supplier Key" needs to be maintained centrally in the gateway only once, and all access permissions are managed uniformly through RBAC and the policy framework. When a new model or supplier is integrated, it can be made available to the entire organization immediately through the same interface.

In AI Gateway, prompts, tools, and agent configurations are all treated as manageable assets. Development teams can define the following configurations in Playground and save them as named templates:

  • System Prompt
  • User Prompt
  • Input Variables
  • MCP Tools
  • Security
  • Model Parameters

Each template supports version control, allowing team members to safely iterate on and optimize configurations without overwriting each other’s changes. If a problem arises with a new version, you can quickly roll back to a previous version.

As a result, AI Gateway effectively provides enterprises with a unified platform for managing prompt and agent configurations. Once a specific configuration has been validated and is ready for deployment, it can be published as an Agent App.

Hongke AI Gateway's security mechanisms operate simultaneously at both the input and output stages, enabling defense-in-depth.

Input Phase Before a request is sent to the model, the system will automatically check for:

  • Personally Identifiable Information (PII)
  • Prompt Injection Attacks
  • Prohibited Content
  • Prohibited Topics

And implement them in accordance with corporate strategy:

  • Interception
  • Desensitization/Masking
  • Replace
  • Content Conversion

Output Phase After the model generates a response, the system performs another security review to check for:

  • Harmful Content
  • Biased Content
  • Hallucination Output (Hallucination)
  • Policy Violations
  • Risk of Sensitive Data Leaks

And based on the strategy, decide:

  • Return the result
  • Editing Results
  • Refuse to Export

AI Gateway can also be integrated with:

  • OpenAI Moderation
  • AWS Guardrails
  • Azure Content Safety
  • Azure PII Detection

It supports existing security services such as [...], while also enabling organizations to create custom rules using profiles or Python code.

Because all security policies are enforced uniformly at the gateway layer, security and compliance teams can implement organization-wide AI usage guidelines in a consistent and auditable manner, making this particularly suitable for highly regulated industries such as healthcare, finance, and insurance.

The Agent App is an agent application built on the AI Gateway that can be delivered directly to end users.

Unlike Playground, the Agent App provides a streamlined and controlled interactive interface. Business teams, testers, or internal users can directly experience how the Agent actually works without having to interact with the underlying prompts, tool configurations, or model parameters.

At the same time:

  • Keep the prompt configuration locked
  • Tool permissions are strictly controlled
  • Security protection remains in effect
  • Centralized Management of Model Routing Policies

This makes the Agent App ideal for:

  • User Acceptance Testing (UAT)
  • Internal Business Pilot Program
  • High-Rise Model Unit
  • Enterprise Copilot
  • Knowledge Q&A Assistant

The product team and the platform team can both provide an open user experience while maintaining full control over configuration and governance rules.

Professional technical support from HONGKEI to help you succeed in your project.

Leveraging decades of technical expertise in automotive electronics, industrial automation, data security, and life sciences, Hongke provides customers with end-to-end solutions ranging from critical components and precision instruments to system-level platforms. We do more than just supply products; we are committed to helping our customers achieve long-term success through professional technical consulting and innovative optimization.

 

Contact Hongke to help you solve your problems.

Let's have a chat