◈ Hub Live Tracker Scan Delta Methodology ↗ Main Site
Research Methodology · v1.0

AI Exposure / Methodology

Published: Feb 24, 2026 Maintained by: Wadsworth Digital Data Source: Shodan.io

This document defines how we detect, classify, and quantify exposed AI infrastructure on the public internet. It covers the detection methodology for each category, the risk framework we apply, known limitations, and what the numbers actually mean.

Contents
  1. What "Exposed" Means
  2. Detection Methodology
  3. Category Definitions & Risk Profiles
  4. Risk Taxonomy
  5. Limitations & Caveats
  6. Ethical Scope

01 What "Exposed" Means

An AI system is considered exposed when it is reachable from the public internet and does not require authentication to access its primary interface or API. This means anyone with an internet connection and a browser — or a simple curl command — can interact with the system directly.

Exposure does not require a vulnerability or a misconfiguration in the traditional sense. Many of these systems are exposed by design: default configurations prioritize developer convenience over network hardening. A system that was spun up for a weekend project, a demo, or a local experiment is indistinguishable from an intentionally public service once it has a routable IP address.

Key distinction: We are counting systems that are accessible without credentials — not systems that are theoretically vulnerable. Every instance in our count can be reached by an anonymous request right now.

The counts represent a point-in-time snapshot of what Shodan's passive scan index has catalogued. They do not represent every exposed instance on the internet — only those Shodan has crawled recently.

02 Detection Methodology

All data is sourced from Shodan.io's passive internet scan index via the Host Count API (/shodan/host/count). Shodan continuously crawls the IPv4 address space, sending lightweight probe packets to indexed ports and recording service banners, HTTP headers, TLS certificates, and product fingerprints.

We do not actively probe any systems. We query Shodan's pre-existing index — the same data available to any Shodan account holder. We never connect to, interact with, or send requests to any discovered host.

Query Reference Table

Category Shodan Query Primary Signal Default Port
Ollama port:11434 product:"Ollama" Product banner match 11434
Jupyter port:8888 product:"Jupyter" Product banner match 8888
Open WebUI "Open WebUI" HTTP title/body match varies
n8n port:5678 "n8n" HTTP body + port 5678
LangChain "langserve" port:8000 LangServe header match 8000
AutoGPT "AutoGPT" port:8080 HTTP title/body + port 8080
Queries use a combination of port and product/banner fingerprint to reduce false positives. A host matching only on port would produce significant noise; the product match confirms the service identity.

Scan Cadence

Scans run automatically once per day at 08:00 ET. Each scan issues one API call per category (six total), records the count, appends to the historical dataset, and redeploys the public tracker. Total scan window: under 60 seconds. Results are published within minutes of scan completion.

03 Category Definitions & Risk Profiles

Each category below represents a distinct class of AI infrastructure. We define what the software is, why it ends up exposed, what the attack surface looks like, and what a bad actor can do with unauthenticated access.

Exposed Ollama Instances
CRITICAL Port 11434

Ollama is an open-source framework for running large language models locally. It exposes a REST API at port 11434 for model inference, management, and streaming completions.

Default configuration binds to 0.0.0.0 (all interfaces) with no authentication. Developers running Ollama on cloud instances or home servers with port forwarding expose it globally without realizing it.

shodan: port:11434 product:"Ollama"
  • Free compute theft — anyone can run inference against any model loaded on the instance, at the host's expense
  • Model enumeration — /api/tags reveals all models installed, including proprietary fine-tunes
  • System prompt extraction — custom Modelfiles and system prompts are retrievable, exposing proprietary configurations
  • Data exfiltration vector — if the Ollama instance is used for RAG, the retrieval context may be accessible through prompt manipulation
  • Lateral movement — the host running Ollama is accessible; the API can sometimes be used to probe the local environment
Unsecured Jupyter Notebooks
HIGH Port 8888

Jupyter Notebook/Lab is an interactive development environment for Python, R, Julia, and other languages. It runs a web server that provides a browser-based code execution interface.

Commonly launched with --ip=0.0.0.0 --no-browser --allow-root flags for remote access on cloud VMs. Token authentication is sometimes disabled for convenience. AI/ML workloads on cloud spot instances are particularly common.

shodan: port:8888 product:"Jupyter"
  • Arbitrary code execution — the Terminal tab provides a full shell on the host system; no exploit needed
  • Training data theft — datasets, model weights, and experiment results stored on the host are fully accessible
  • Credential harvesting — .env files, API keys in notebooks, and cloud provider credentials are trivially accessible
  • Cryptomining implant — GPU-equipped instances are frequently targeted for cryptomining via notebook code execution
  • Supply chain injection — notebooks used in CI/ML pipelines can be backdoored to inject malicious behavior into production models
Open WebUI Interfaces
CRITICAL Port varies

Open WebUI (formerly Ollama WebUI) is a self-hosted frontend for Ollama and other LLM backends. It provides a ChatGPT-style interface with conversation history, system prompts, model switching, and multi-user support.

Often deployed via Docker with -p 3000:8080 without a reverse proxy or auth layer. User registration is frequently set to "open" by default, allowing anyone to create an account and access the LLM backend.

shodan: "Open WebUI"
  • Unrestricted LLM access — full model inference at the host's compute and API cost
  • Conversation history access — other users' conversation histories may be accessible depending on configuration
  • System prompt theft — custom personas, RAG configurations, and business logic embedded in system prompts are exposed
  • Admin panel access — default admin credentials or open registration may grant full administrative control
  • Backend pivoting — Open WebUI often connects to Ollama or OpenAI; exposed instances can be used to exhaust upstream API quotas
Open n8n Automation Servers
HIGH Port 5678

n8n is an open-source workflow automation platform. It connects APIs, services, and databases through visual workflows — similar to Zapier or Make, but self-hosted. Increasingly used as the orchestration layer in AI agent pipelines.

Default n8n installations run without authentication on port 5678. Instances are often spun up quickly for prototyping AI agent flows and left accessible, particularly in the AI agent community where n8n is a popular building block.

shodan: port:5678 "n8n"
  • Credential exposure — workflow nodes store API keys, OAuth tokens, database passwords, and webhook secrets in plaintext
  • Workflow execution — an attacker can trigger existing workflows, potentially sending emails, posting to APIs, or modifying databases
  • Business logic theft — automation workflows encode proprietary business processes and integrations
  • Webhook manipulation — exposed webhook trigger nodes can be replayed or spoofed to inject malicious data into business pipelines
  • AI agent hijacking — n8n is commonly the orchestration layer for AI agents; control of n8n = control of the agent's actions
LangChain API Endpoints
MEDIUM Port 8000

LangChain is a framework for building LLM-powered applications. LangServe is its deployment layer, which exposes chains and agents as REST endpoints. Detected via the x-langchain-* response header fingerprint.

LangServe defaults to no authentication. Developers exposing chains for internal use or demo purposes often skip the auth layer entirely. The count in this category is declining as LangServe's market position shifts and teams adopt more production-hardened deployment patterns.

shodan: "langserve" port:8000
  • Prompt injection — exposed chains can be manipulated via crafted inputs to bypass intended behavior or extract system prompts
  • Proprietary chain theft — chain definitions expose application architecture and business logic
  • Tool abuse — chains with tool-calling enabled (web search, database queries, code execution) can be misused by unauthenticated callers
  • API cost exhaustion — every request to an exposed LangServe endpoint consumes the host's upstream LLM API quota
AutoGPT Instances
MEDIUM Port 8080

AutoGPT is an autonomous AI agent framework that enables goal-directed task execution. It can browse the web, write and execute code, manage files, and interact with external services autonomously.

Experimental deployments, research instances, and demos. AutoGPT's count has declined significantly as the project matured and users shifted to newer agent frameworks. Remaining instances are likely abandoned or unmonitored.

shodan: "AutoGPT" port:8080
  • Autonomous agent hijacking — an attacker can task the agent with arbitrary goals including data exfiltration, account creation, or external service abuse
  • API key exposure — AutoGPT configurations contain OpenAI and other API keys which are often readable through the web interface
  • Unmonitored execution — abandoned instances may continue running tasks without any human oversight

04 Risk Taxonomy

We apply a three-tier risk classification to each category based on the severity of unauthenticated access and the breadth of potential harm.

LevelDefinitionExamples
CRITICAL Unauthenticated access enables direct code execution, full system compromise, or complete control over AI workloads and data. No exploit required — the access is the vulnerability. Ollama (compute theft + data exfil), Open WebUI (full LLM access + conversation history)
HIGH Unauthenticated access enables significant data exposure, credential theft, or execution of consequential business logic. Exploitation requires minimal effort. Jupyter (code execution via Terminal), n8n (credential exposure + workflow execution)
MEDIUM Unauthenticated access exposes proprietary logic, enables API cost exhaustion, or allows manipulation of AI system behavior. Impact depends on what the system is connected to. LangChain (prompt injection + tool abuse), AutoGPT (agent tasking)

05 Limitations & Caveats

Passive Scan Coverage Gap

Shodan does not have real-time coverage of the entire IPv4 address space. Some hosts are crawled infrequently; newly exposed instances may not appear for hours or days. Our counts represent a lower bound on actual exposure — the real number is likely higher.

False Positives

Product fingerprint queries reduce but do not eliminate false positives. A service running on port 11434 that returns a banner matching the Ollama product string but is not actually Ollama would be included in the count. In practice, fingerprint false positives for these categories are rare.

Shodan Re-Index Events

Shodan periodically performs bulk re-crawls of IP ranges that have not been indexed recently. This can cause sudden spikes in counts that reflect newly catalogued hosts, not newly exposed ones. Hosts may have been exposed for weeks before appearing in our data. We flag data points where the daily delta exceeds 2,000% as anomalous and mark them for review.

Example: On Feb 24, 2026, Jupyter Notebook counts spiked from 111 to 5,415 — a 4,779% increase in 24 hours. This was flagged as a probable Shodan re-index event and marked ⚑ on the tracker. The count is directionally significant but the delta does not represent 5,304 systems becoming newly exposed overnight.

Authentication State

Our queries detect that a service is running and publicly reachable. We do not verify authentication state beyond what is visible in Shodan's indexed banner. Some instances in our count may have authentication enabled at a layer not visible to Shodan's passive probe (e.g., a reverse proxy in front of the service). Our count is therefore a ceiling on unauthenticated exposure for some categories.

No Raw IP Data Published

We publish aggregate counts only. We do not publish, retain, or distribute lists of IP addresses, hostnames, or any device-identifying information from Shodan results.

06 Ethical Scope

This research is conducted within the following ethical boundaries:

The goal is simple: the people building and deploying AI infrastructure should know what the exposure landscape looks like. Awareness is the first step toward hardening.

If you believe your system is included in this count and want to understand how to harden it, run a free security scan or contact Wadsworth Digital for a full assessment.