Corelight Bright Ideas Blog: NDR & Threat Hunting Blog

Shadow AI Detection: See and Govern AI Traffic | Corelight

Written by Tim Chiu | Jul 22, 2026 2:30:00 PM

AI adoption inside the enterprise didn't ask for permission. It arrived through browser tabs, code editors, and meeting transcription bots, quietly stitching itself into daily workflows long before security teams could write policy around it. The result is a familiar story with a new villain, a sprawling, unmanaged attack surface that lives in your network traffic but nowhere in your asset inventory.

We call it shadow AI, and it's the blind spot you didn't plan for or budget for.

With our latest Sensor release, v29.1.2, Corelight closes that gap. This major expansion to our shadow AI application identification adds detection for over 80 new AI services through network log enrichments, mapped across 100-plus domain-pattern rules. The result is clean, evidence-based visibility into who is talking to which AI service, surfaced directly in the Zeek logs you already trust.

The new blind spot is hiding in plain sight

Security teams already know the drill with shadow IT. Shadow AI is the same problem wearing a smarter hat.

An employee pastes a sensitive customer record into a free LLM. A developer pipes proprietary source code through an AI coding assistant to speed things up. A team adopts a meeting transcription tool that quietly ships audio and transcripts to a third party. None of it shows up in your sanctioned tool list, yet all of it generates network traffic the moment it touches an external endpoint.

That last detail is the one that matters. Endpoint Detection and Response (EDR) has visibility only into devices it’s installed on, and acceptable-use policies can't be enforced when your tools aren’t present. The network, however, sees everything. If a service phones home, transfers data, or moves laterally in your network, your sensor knows.

This is the core truth of governance: You can't govern what you can't see. Shadow AI detection turns invisible traffic into a structured, searchable inventory you can actually act on.

What the new release actually delivers

Shadow AI application identification enriches network logs to surface over 80 new distinct AI services, mapped across more than 100 domain-pattern rules. In today’s AI landscape, a single provider can operate across multiple domains. Corelight collapses that into a single, unified service entry, so your inventory reflects the actual landscape of AI usage, not the underlying domain sprawl.

Coverage spans 13 important categories, including Commercial LLM APIs, Chinese AI providers, AI proxies and aggregators, AI coding assistants, and AI agents. These are some of the riskiest AI categories that fall into the realm of shadow AI.

Because this detection is powered by Corelight's evidence-first enrichment, the output isn't a standalone alert with no explanation. It includes context: Structured metadata you can pivot on, baseline against, and feed straight into your SIEM or SOAR workflows.

Where the real risk lives

Not all AI traffic carries the same weight. Some of it is valid and used as a productivity tool. Some of it is a data exfiltration incident in slow motion. Knowing the difference is what separates inventory from intelligence, which is why we've organized the risk into tiers.

High risk: Data sovereignty and foreign infrastructure

A significant portion of the newly detected AI services routes data through foreign-controlled infrastructure, raising direct data sovereignty and exfiltration concerns. For regulated industries and government-adjacent organizations, this is the point where shadow AI stops being a productivity conversation and starts becoming a compliance issue. Detection in this category helps security teams identify when enterprise data is leaving for jurisdictions outside their control, flag that activity for legal and compliance review, and build a defensible record of unsanctioned access. Visibility in this category is non-negotiable.

High risk: Evasion through AI proxies and aggregators

AI proxy and aggregator services are among the most sophisticated evasion categories in the shadow AI landscape. These services act as routing layers between users and the actual AI models they reach, intentionally abstracting the downstream endpoint. For security teams, that abstraction is the problem. An employee's traffic may appear as a single, low-risk outbound connection while silently funneling requests to any number of unsanctioned models behind it. Detecting proxy and aggregator usage exposes this layer of indirection, giving analysts the ability to identify policy evasion at the infrastructure level rather than chasing individual model endpoints that shift or expand over time.

Medium risk: Intellectual property exposure through AI coding tools

AI coding assistants sit at a uniquely sensitive intersection: They are deeply embedded in developer workflows and directly handle some of an organization's most valuable data. Every context window shipped to a cloud-based coding service may contain proprietary source code, internal architecture patterns, authentication logic, or credentials. Detecting this category helps security teams confirm which external coding tools are in active use, identify where intellectual property may be flowing without data loss prevention (DLP) oversight, and surface unsanctioned developer tooling before it becomes an exposure event. For organizations where source code is a core competitive asset, visibility into this category is as important as any data classification policy.

Medium risk: Autonomous AI actions and agentic tooling

Agentic frameworks represent one of the more complex detection challenges in the shadow AI landscape. Unlike single-turn LLM queries, agentic tools execute automated, multi-step workflows that chain together actions, make decisions, and interact with external systems, often with minimal human oversight. That autonomy makes them powerful for legitimate automation, but also attractive for attack orchestration and difficult to attribute when something goes wrong. Detecting this category of traffic gives security teams a critical signal: Not just that an AI tool is in use, but that autonomous, goal-directed activity may be occurring on their network. Knowing when and where agentic workflows fire is the first step to distinguishing sanctioned automation from something far less benign.

The remaining categories complete a comprehensive picture of the enterprise AI landscape. Commercial LLM APIs establish the broadest baseline for shadow AI inventory, capturing the day-to-day sprawl of unsanctioned model usage across the organization. AI writing and meeting tools surface quieter but still consequential risks, including Sensitive business content processed outside DLP controls as well as audio, transcripts, and meeting summaries sent to third-party services without IT review. Model hubs, ML platforms, and vector databases signal active experimentation and internal AI development that may be operating well outside governance frameworks. Enterprise AI and specialty platforms reveal structured, often departmental adoption patterns, giving security teams an early view of where AI initiatives are taking root across the business.

From visibility to enforcement

Detection is the foundation. Once shadow AI traffic shows up as structured evidence, your team can put it to work:

  • Build a living shadow AI inventory across the enterprise, with foreign providers flagged for data sovereignty review
  • Identify data exfiltration risk when sensitive content flows to LLM APIs and coding assistants outside your DLP controls
  • Spot threat actor tooling, including agentic frameworks used to automate attacks
  • Baseline volume and behavior to catch abnormal usage before it becomes an incident
  • Feed acceptable-use policy violations directly into your SOAR and SIEM for enforcement

This is the practical payoff for SOC analysts, detection engineers, SecOps leaders, and CISOs alike: A defensible answer to the question every board is starting to ask, "What AI are our people actually using, and where is our data going?"

Get started with Sensor v29.1.2

Shadow AI isn't going to slow down, and frankly, neither should your visibility into it. Sensor v29.1.2 gives you the ground-truth evidence to inventory AI usage, prioritize the riskiest traffic, and enforce policy with confidence, all built on the open, evidence-first foundation you already rely on.

The AI you can't see is still your responsibility. Now, with Corelight, you can see it. Upgrade to Sensor v29.1.2 and bring shadow AI into the light. If you’re not a current customer, learn more about Corelight today.