AI Agents

Clickport shows you which AI tools read your site: ChatGPT, Claude, Perplexity, Gemini and others. For every engine you see what it read, what it used in live answers, how many visitors it sent you, and whether they converted. And for every page, you see what AI reads next to what humans read.

Real behavior, not estimates. Every number comes from requests your server actually received and sessions your site actually had. Clickport does not run prompts at AI models to guess your visibility, and there is no synthetic score. If ChatGPT read your pricing page 40 times last week, that is what you see.

Why this needs a connector

AI tools do not load the tracking script, so normal analytics never sees them. The only place their visits exist is your own server. A small connector passes that information on to Clickport: which page was requested, when, and by whom. Clickport does the rest and works out which requests came from AI.

The connector itself knows nothing about bots. When new AI agents appear, Clickport recognizes them automatically and your connector never needs an update.

Setup takes a few minutes and has two steps: create a key in Settings → Integrations → AI Agents, then follow the guide below that matches your site.

Three kinds of AI reads

Every AI hit on your site is one of three things:

  • Live retrieval. An assistant fetches your page right now, in the middle of a real person's conversation, because it needs your content for the answer. This is the closest measurable thing to appearing in an AI answer.
  • Search indexing. A crawler collects your page for an AI search index, so the engine can find and cite it later.
  • Model training. A crawler collects your page to train future models. You get nothing back today. At best, tomorrow's model knows you exist.

In the dashboard, live retrievals appear as Cited. Indexing and training crawls together appear as Ingested. For comparison, Clickport also records classic search crawlers like Googlebot and Bingbot, but they never count into any AI number.

What you see in the dashboard

Sources → AI: the engine funnel

Each AI company is a row with four stages:

  • Cited: pages fetched live during a real person's AI conversation. When ChatGPT answers someone and pulls your page in to do it, that is one citation. This is the closest measurable thing to "my site appeared in an AI answer".
  • Ingested: pages crawled for AI search indexes and model training (for example GPTBot or ClaudeBot collecting training data).
  • Visited: human visits from the AI Search channel, the people who clicked through from an AI answer.
  • Converted: those AI-referred visits that completed one of your goals.

The stages tell each engine's story. Some engines cite thousands of times without crawling ahead: they fetch your page on demand at answer time. Others ingest heavily for training and never cite. Some send visitors without ever fetching, because their answers ride a search index crawled by someone else. The funnel makes those strategies visible for your own site.

SourcesLocationsTechnologiesCampaigns
ChannelsSourcesURLsAI
3061 citations · 10.9k crawls · 497 AI visits · 55 conversions
Engine
Cited ↓
Ingested
Visited
Converted
OpenAI1364482021224
Perplexity1027298817419
Anthropic3861743587
Google204962385
Microsoft66310120
Mistral149830
Cross filters don't apply to agent traffic.

Pages → AI: what AI reads, next to what humans read

Every page with agent activity is a row with Cited (live-answer fetches), Ingested (crawls for AI indexes and model training), and Visitors (unique human visitors of the same page in the same range). The little caret unfolds a per-engine split for that page, including how many visitors each engine's answers sent to it.

Two icons on each row connect the AI world to the rest of your dashboard:

  • The bot icon draws that page's citations as a dashed line on the main chart, on the same scale as your visitors. Click again to remove it. Engine rows in Sources → AI do the same on click.
  • The human icon applies a regular page filter, exactly as if you had clicked the page in Top Pages, so the whole dashboard shows that page's human traffic. Both icons can be active at once, which puts a page's human visits and citations side by side on one chart.
PagesSessionsGoalsJourneys
TopEntryExitSearch404AI
Page
Cited ↓
Ingested
Visitors
/blog/google-analytics-alternative486718892
EngineCitedIngestedAI visitors
OpenAI21427538
Perplexity17623961
Anthropic9620417
/docs/getting-started342519573
/pricing2913531108
/blog/cookieless-tracking227311402
Cross filters don't apply to agent traffic.

The channels list

Below your traffic channels, a separate row shows total AI agent fetches for the selected range. It is deliberately not a channel: agent fetches are not visits and never count toward your visitor numbers or percentages. Clicking the row opens the AI view.

SourcesLocationsTechnologiesCampaigns
ChannelsSourcesURLsAI
Channel
Visitors ↓
%
Organic Search5,20443%
Direct3,02425%
Organic Social1,42812%
AI Search9848%
Email4924%
AI Agents14.0k fetches

Bot Center: verification

Every agent hit is verified against the operator's published IP ranges, where the operator publishes them:

  • Verified: the request came from the operator's published address space.
  • Spoofed: the user agent claims a known bot, but the IP is outside the operator's published ranges. Scrapers love pretending to be well-known crawlers; Clickport catches them.
  • Unverifiable: the operator publishes no ranges to check against.

Spoofed hits are excluded from every AI number in the dashboard. The per-agent verification detail lives in Settings → Bot Center.

Six operators publish ranges today: OpenAI, Anthropic, Perplexity, Google, Microsoft and Apple. Everyone else shows as unverifiable, which only means there is nothing to check against, not that something is wrong. One special case: genuine Claude-User fetches arrive from user-side addresses rather than Anthropic's servers, so they always show as unverifiable.

AI agent verification
GPTBot OpenAI3,204 verified118 spoofed
ChatGPT-User OpenAI1,616 verified
ClaudeBot Anthropic1,286 verified41 spoofed
PerplexityBot Perplexity964 verified210 spoofed
Claude-User Anthropic457 verified
CCBot Common Crawl152 unverifiable
Verified = hits from the operator's published IP ranges. Spoofed hits are excluded from every AI number in the dashboard.

Who's who: the agents we recognize

Clickport recognizes around 60 AI agents and sorts each one into the three read types automatically. These are the ones you will actually meet:

OpenAI

  • ChatGPT-User fetches pages live while ChatGPT answers someone. Cited.
  • OAI-SearchBot crawls for ChatGPT's search index. Ingested.
  • GPTBot collects training data. Ingested.

Anthropic

  • Claude-User fetches pages live during Claude conversations. Cited.
  • Claude-SearchBot crawls for Claude's search index. Ingested.
  • ClaudeBot collects training data. Ingested.

Perplexity

  • Perplexity-User fetches pages live at answer time. Cited.
  • PerplexityBot crawls for Perplexity's index. Ingested.

Google

  • Gemini-Deep-Research, Google-NotebookLM and Google-Agent fetch live for Gemini's research and agent tools. Cited.
  • Google-CloudVertexBot crawls for companies building AI search on Google's cloud. Ingested.
  • Google-Extended and GoogleOther collect training data. Ingested.

Note what is missing: AI Overviews have no crawler of their own. They work from the normal Googlebot index.

Microsoft

There is no Copilot crawler at all. Copilot answers from the Bing index, so plain Bingbot (listed under Classic Search) is quietly doing the AI work too.

Meta

meta-externalagent and FacebookBot collect training data, meta-webindexer builds an index, and meta-externalfetcher can fetch live for Meta AI. Ingested, except the fetcher, which counts as Cited.

ByteDance

Bytespider, TikTokSpider, Doubaobot and imageSpider all collect training data. ByteDance runs large AI assistants, but discloses no agent that fetches live or sends anyone back. Ingested, always.

Apple

Applebot indexes for Siri, Spotlight and Apple Intelligence. Applebot-Extended is the separate training crawler. Both Ingested.

Amazon

Amazonbot and Amzn-SearchBot index for Alexa and Rufus (Ingested). Amzn-User, NovaAct and AmazonBuyForMe fetch live, including shopping agents browsing on a customer's behalf (Cited).

Everyone else

Clickport also recognizes agents from Mistral, Cohere, DuckDuckGo, Kagi, Brave, You.com, Phind, Andi, Linkup, Huawei, xAI, DeepSeek, Alibaba, Zhipu AI, Moonshot AI, Baidu, Yandex, AI2, Diffbot, QuillBot and Cognition's Devin, plus the research crawlers CCBot (Common Crawl) and ClueWeb-Crawler (Carnegie Mellon), whose datasets feed many labs' training runs. New agents appear monthly. Clickport adds them centrally, and your connector never needs an update.

Every bold name above is also the exact token for your robots.txt. Most operators separate training from search, so you can block training crawls without disappearing from AI search.

How the engines behave

The funnel looks different for every engine, and that is the point. A few patterns you will recognize in your own data:

  • Fetch-at-answer-time engines. ChatGPT and Perplexity fetch pages live while answering. Heavy Cited counts, and the reads that most often turn into visits.
  • Index-first engines. Copilot and Google's AI Overviews answer from an existing search index. They can cite you and send visitors without a single AI agent in your logs. Zero crawls does not mean zero AI presence.
  • Walled gardens. Meta's and ByteDance's assistants live inside their own apps. On the open web their crawlers mostly collect training data, and reads rarely turn into visits.
  • Training-only crawlers. CCBot, ClueWeb and friends take content for tomorrow's models and send nothing today.

Connect with a Cloudflare Worker

The best choice when your site runs through Cloudflare. It never slows your site down, and it sees every visit, even ones served from cache.

  1. In Clickport, open Settings → Integrations → AI Agents and create your key. The Worker code shown there already contains it.
  2. In the Cloudflare dashboard, create a Worker and paste the code.
  3. Give the Worker the route yourdomain.com/* so it covers your whole site.
  4. Optional but recommended: store the key as a Worker secret named CLICKPORT_KEY instead of leaving it in the code.

That is it. Your data shows up in Sources → AI within a minute of the first AI visit.

Connect with the WordPress plugin

The easiest choice when your site runs on WordPress. You need the Clickport plugin, version 1.2.0 or newer. No code needed.

  1. Install or update the plugin from the WordPress guide.
  2. In Clickport, open Settings → Integrations → AI Agents and create your key.
  3. In WordPress, open Settings → Clickport and paste the key into the AI Agents connector field.
Caching plugins hide some visits. Some caching plugins and CDNs answer a request before WordPress wakes up, and those visits cannot be counted. The numbers you get are real, just incomplete on heavily cached sites. If your site also runs through Cloudflare, use the Worker connector instead: it sees everything.

Connect with the Node.js package

For Next.js, Express, and plain Node servers. Works on Vercel (all plans, including Hobby) and self-hosted. On Netlify, use the Edge Function connector in the next section instead: it covers every framework there, Next.js included. Delivery rides the platform's own after-response hook, so it adds zero latency, and no report is lost when a serverless environment freezes after responding. That last part matters: a plain fire-and-forget request from middleware is silently dropped on these platforms.

  1. Install the package: npm install @clickport/agents
  2. In Clickport, open Settings → Integrations → AI Agents and create your key. Set it as an environment variable named CLICKPORT_AGENT_KEY wherever your app runs.
  3. Wire in the adapter for your framework.

Next.js: create this file at the project root. On Next 12.2 to 15 it is called middleware.ts; Next 16 renamed the convention to proxy.ts. Same content either way:

import { withClickportAgents } from '@clickport/agents/next';

export default withClickportAgents();

export const config = {
  matcher: ['/((?!_next/static|_next/image|favicon.ico).*)'],
};

Already have middleware? Pass it through: withClickportAgents(myMiddleware). One honest limitation: middleware runs before your routes, so these reports carry no response status. A request that ends up a 404 still counts as a fetch. For exact status codes use the Express adapter below, or wrap individual route handlers with withAgentTelemetry from the same package.

Express: one line, full fidelity (real status codes and durations, batched delivery):

import { clickportAgents } from '@clickport/agents/express';

app.use(clickportAgents());

Plain node:http: instrument(server) from @clickport/agents/node attaches the same reporting to an existing server without touching your handlers.

The package is MIT licensed, has zero dependencies, and forwards a fixed whitelist of request metadata: never cookies, never authorization headers, never request bodies.

Connect on Netlify

For any site hosted on Netlify, whatever it is built with: static, Astro, Hugo, or Next.js. One Edge Function file sees every request, including cached ones, and reports after the response is already on its way. One connector per site is enough.

  1. Save the code below as netlify/edge-functions/clickport-agents.js in your project.
  2. In Netlify, open Site configuration → Environment variables and add CLICKPORT_AGENT_KEY with the key from Settings → Integrations → AI Agents. Never paste the key into the file itself.
  3. Deploy, then click Test my connector in the same Clickport settings section.
// netlify/edge-functions/clickport-agents.js
const CLICKPORT_ENDPOINT = 'https://clickport.io/api/agent-visits';

// Static assets are skipped: agent analytics cares about content, not styling.
const ASSET_RE = /\.(css|js|mjs|png|jpe?g|gif|svg|ico|woff2?|ttf|otf|eot|webp|avif|mp4|webm|mp3|map|zip|gz|txt)$/i;

// Request headers forwarded verbatim; Clickport's servers decide trust.
// Never cookies, never authorization.
const HEADER_WHITELIST = [
  'user-agent', 'referer', 'accept-language',
  'sec-ch-ua', 'sec-ch-ua-mobile', 'sec-ch-ua-platform',
  'signature', 'signature-input', 'signature-agent',
  'x-forwarded-for', 'x-nf-client-connection-ip',
];

export default async function clickportAgents(request, context) {
  const started = Date.now();
  const response = await context.next();

  try {
    const method = request.method;
    if (method !== 'GET' && method !== 'HEAD') return response;
    const url = new URL(request.url);
    if (ASSET_RE.test(url.pathname)) return response;
    if (response.status < 200 || response.status >= 300) return response;

    const key = globalThis.Netlify?.env?.get?.('CLICKPORT_AGENT_KEY')
      ?? globalThis.Deno?.env?.get?.('CLICKPORT_AGENT_KEY');
    if (!key) return response;

    const headers = {};
    for (const name of HEADER_WHITELIST) {
      const value = request.headers.get(name);
      if (value) headers[name] = value.slice(0, 2048);
    }
    // context.ip is the connection peer Netlify saw.
    if (context.ip) headers['remote-addr'] = context.ip;

    const line = JSON.stringify({
      ts: new Date(started).toISOString(),
      path: url.pathname,
      method,
      status: response.status,
      duration_ms: Date.now() - started,
      headers,
      connector: 'netlify',
      connector_version: 'edge-1',
    });

    const send = fetch(CLICKPORT_ENDPOINT, {
      method: 'POST',
      headers: {
        'Content-Type': 'application/x-ndjson',
        'Authorization': `Bearer ${key}`,
      },
      body: line + '\n',
    }).catch(() => {});

    if (typeof context.waitUntil === 'function') context.waitUntil(send);
  } catch {
    // Reporting must never break the site.
  }

  return response;
}

export const config = {
  path: '/*',
  // Common static files never invoke the function at all.
  excludedPath: ['/*.css', '/*.js', '/*.png', '/*.jpg', '/*.jpeg', '/*.svg', '/*.ico', '/*.woff', '/*.woff2', '/*.webp', '/*.avif'],
};

Connect with the nginx log shipper

For self-hosted servers behind nginx. This is the most accurate connector: it reads the access log, so no cache layer can hide a visit. It is also the only one with history: on install it can backfill up to 35 days of rotated logs, so you see real agent data immediately instead of waiting for new visits.

  1. Download the shipper: a single file, zero dependencies, Node.js 20 or newer.
  2. Create your key in Settings → Integrations → AI Agents.
  3. Backfill once, then start live shipping:
CLICKPORT_AGENT_KEY=ck_... node clickport-agent-shipper.mjs --backfill
CLICKPORT_AGENT_KEY=ck_... node clickport-agent-shipper.mjs

Run the live shipper under systemd or pm2; the README in the download contains ready-made service files. The standard combined log format works as-is. Configuration is three optional environment variables: NGINX_ACCESS_LOG for a custom log path, HOST_FILTER when one log file mixes several vhosts, and SKIP_PATHS for path prefixes you never want reported.

Run the backfill exactly once. Rerunning it ships the same log lines again and duplicates rows. Backfill first, then start the live shipper; the two do not overlap.

Connect from any server

For developers: any backend, edge function, or log pipeline can report visits. Send one JSON line per request your server saw:

POST https://clickport.io/api/agent-visits
Authorization: Bearer ck_your_connector_key
Content-Type: application/x-ndjson

{"ts":"2026-07-13T09:00:00Z","path":"/blog/post","method":"GET","status":200,"duration_ms":42,"headers":{"User-Agent":"...","Remote-Addr":"203.0.113.7"},"connector":"rest"}

Send every content request. Clickport keeps the AI agents and discards the human hits. Up to 5,000 lines per request, gzip supported, and the endpoint responds 204 on success. The full field reference is in the API documentation.

Reading the numbers honestly

  • AI visits are a floor. Assistants often strip referrers, so some AI-driven visits land in your Direct channel and cannot be attributed to an engine. The Visited column undercounts rather than guesses.
  • A citation is a fetch, not a guaranteed appearance. An engine may fetch a page while composing an answer and use only part of it, or fetch it more than once for one answer. It is the strongest signal available from real behavior, not a screenshot of the answer.
  • Collection starts at connection. Requests that happened before the connector was installed were never recorded anywhere Clickport can see, so there is no retroactive history.
  • Cross filters do not apply to agent traffic. Agent hits carry no country, source, or device, so the AI views ignore the dashboard's filters. The human numbers inside them (Visitors, Visited, Converted) are real session data and pair with the rest of your dashboard through the icons described above.
  • Classic search crawlers are not AI. Googlebot and Bingbot visits are collected and visible in the Bot Center, but they never count into citations or crawls.
  • Agent hits never touch your bill. They do not count as pageviews and do not consume your plan's quota.

Privacy and data handling

  • Connectors send request metadata only: path (query strings are stripped), method, status, timing, user agent, and the requester's IP for verification.
  • The IP is checked against published operator ranges at ingest and immediately discarded. Only the verdict (verified, spoofed, unverifiable) is stored.
  • Agent data is retained for 13 months.
  • None of this involves your human visitors: their analytics continue to come from the cookie-free tracker alone.

Troubleshooting

The status says "Not connected" after setup

The status turns green on the first reported AI visit, not when you create the key. Small sites can go hours between AI visits, so give it time. To check your setup right away, click Test my connector in Settings → Integrations → AI Agents: Clickport visits your homepage with a test user agent, and a working connector reports it back within seconds. Test hits are never counted in your dashboard. The same test works by hand from any machine: curl -A "ClickportConnectorTest/1.0 (+https://clickport.io/docs/ai-agents)" https://yourdomain.com/.

Rows appear, but everything is "unverifiable"

Your connector is not passing the requester's IP. In the Worker this is automatic. For the API, include the client address as Remote-Addr (or X-Forwarded-For) inside each line's headers object.

WordPress shows fewer reads than expected

See the page-cache warning above: requests served entirely from cache never reach the plugin. The counts you do get are real, just incomplete on heavily cached sites.