Blog Zscaler
Recevez les dernières mises à jour du blog de Zscaler dans votre boîte de réception
Zero Trust Connectivity for Private AI Apps/Models: Who Can Actually Reach Them?
The enterprise AI landscape is shifting fast. While public discourse has centered on ChatGPT, Copilot, Grok, Mythos, and the race to adopt SaaS AI tools, a quieter but equally significant trend is underway: organizations are deploying private AI models inside their own data centers. Whether it's a fine-tuned model trained on internal documents, a private inference server running at the edge, or an agentic AI pipeline connected to proprietary knowledge bases, enterprises are betting that keeping AI on-premises solves their biggest concerns about data sovereignty, regulatory compliance, and sensitive IP protection.
They're right that self-hosting eliminates many of the third-party data risks inherent in cloud AI. But in solving one problem, many organizations are inadvertently creating another: they're securing the model itself, while leaving access to it dangerously wide open.
VPN Was Not Built For This
Here's a scenario that plays out more often than most security teams would like to admit. A team of data scientists or developers needs access to an internal AI inference endpoint. IT provisions VPN access to that network segment, the team connects, and work begins. Simple, fast, done.
Except that VPN access doesn't grant access to one resource. It grants access to everything on that network segment. The AI inference server, adjacent application servers, internal APIs, file shares: all of it becomes reachable the moment a user authenticates to the VPN. The model may be locked down, but the blast radius of a compromised credential, a misconfigured client, or insider misuse is enormous.
This is the fundamental flaw in the “VPN to on-prem AI” model: it conflates network access with application access. A VPN authenticates a device and drops it onto a segment of the corporate network. It knows nothing about whether that specific user should be interacting with a particular AI model or any of the internal resources connected to it. In an AI context, those questions are exactly the ones that matter.
The risks compound quickly. Shadow AI infrastructure is a growing concern: developers spin up self-hosted models on unmanaged GPU workstations, teams deploy unauthorized MCP servers to connect AI agents to internal tools, and rogue AI applications proliferate without security review. These environments often run with no authentication at all, discoverable by anyone already on the network.
Even properly managed AI servers can become data exfiltration tools: a malicious actor, or simply an overly curious insider, can use natural language to extract and summarize sensitive information from across a flat network segment that the VPN trusts implicitly. Traditional DLP tools aren’t built to catch that. The VPN logs tell you who connected. They tell you nothing about what the model was instructed to do.
Access to AI Is an Identity Problem
Securing on-premises AI infrastructure is not fundamentally a network problem. It's an identity problem. The core question isn’t “Is this device on the corporate network?” It's “Is this person, with this device posture, authorized to access this specific AI resource right now, and only this resource?”
That shift in framing changes everything. If identity is the control plane, then access to every component of your AI environment should be governed at the application level, by policy tied to who the user is. A data scientist in the research team should be able to reach the models scoped to their work. A finance analyst should reach only the AI resources relevant to their role. Neither should see the other’s environment, and neither should be able to traverse to adjacent infrastructure they have no business touching.
This is zero trust applied to AI: least-privilege access, identity-verified at every connection, with no implicit trust granted simply because a user is on the network.
How Zscaler Approaches This
Zscaler Private Access (ZPA) is purpose-built for exactly this problem, and its architecture maps directly onto the requirements of securing on-premises AI infrastructure.
ZPA operates on a “segment of one.” Rather than placing a user on a network and trusting them to behave, ZPA brokers a direct, encrypted connection between the authenticated user and the specific application they are authorized to reach, and nothing else. The internal AI environment is never exposed to the internet, and it is never reachable simply because someone has network access. Lateral movement becomes structurally impossible: a user connected to an AI inference endpoint cannot see, scan, or reach anything else on the segment.
The mechanics matter here. A lightweight App Connector deployed next to your AI infrastructure makes only outbound connections to the Zscaler Zero Trust Exchange. There are no inbound open ports. The AI environment is invisible to the public internet and unreachable from the broader corporate network without explicit, identity-based authorization. A compromised credential doesn’t become a skeleton key.
At the moment of access, ZPA validates both user identity (through your existing identity provider such as Okta or Azure AD), and device posture. Is this a managed device? Is the OS patched? Is the endpoint agent running?
ZPA also brings visibility to AI apps that security teams may not even know exist. Using ZPA’s application-type classification capabilities, organizations can automatically identify AI applications running in their environment without requiring traffic decryption. Those apps can then be tagged as sanctioned or unsanctioned.
Unsanctioned AI applications, including unauthorized MCP servers or shadow AI deployments, can be blocked immediately while app owners are engaged to bring them into compliance or decommission them. For sanctioned applications, only after identity and device posture checks pass does a connection get brokered, scoped only to what that user is permitted to reach.

But secure connectivity and Zero Trust Access is not the full the picture. Once a user or an AI agent is connected, what are they actually sending to the model? Zscaler AI Guard extends protection into the session itself, inspecting prompts and responses in real time to block prompt injection attacks, prevent sensitive data from being submitted to the model, and stop confidential information from leaking in model outputs. Combine that with ZPA’s segmentation capabilities and Zscaler’s deception technology, which can lure and expose unauthorized access attempts with decoy resources, and the result is a layered security story that covers discovery, access control, content inspection, and threat detection in one platform.
Together, this means a complete audit trail of who connected and when, what the AI model was asked, and what it returned. In a world where regulatory frameworks are beginning to demand documented access controls and usage logs for AI systems, that auditability is quickly becoming a compliance requirement.
The Bottom Line
Running AI on-premises is the right choice for organizations that need to keep sensitive data, proprietary models, and confidential IP firmly within their own control, out of the hands of third-party AI providers. But the security gains from self-hosting are quickly negated if access to that AI infrastructure is governed by the same broad, implicit-trust network access model that zero trust was designed to replace.
Identity-based, application-level access, enforced by architecture rather than policy documents, is how you close that gap. Zscaler makes this possible.
Cet article a-t-il été utile ?
Clause de non-responsabilité : Cet article de blog a été créé par Zscaler à des fins d’information uniquement et est fourni « en l’état » sans aucune garantie d’exactitude, d’exhaustivité ou de fiabilité. Zscaler n’assume aucune responsabilité pour toute erreur ou omission ou pour toute action prise sur la base des informations fournies. Tous les sites Web ou ressources de tiers liés à cet article de blog sont fournis pour des raisons de commodité uniquement, et Zscaler n’est pas responsable de leur contenu ni de leurs pratiques. Tout le contenu peut être modifié sans préavis. En accédant à ce blog, vous acceptez ces conditions et reconnaissez qu’il est de votre responsabilité de vérifier et d’utiliser les informations en fonction de vos besoins.
Recevez les dernières mises à jour du blog de Zscaler dans votre boîte de réception
En envoyant le formulaire, vous acceptez notre politique de confidentialité.


