ai-seo

v2026.09.24

Optimize content to rank in AI search engines (AI Overviews, Perplexity, ChatGPT) via generative engine optimization (GEO), citability audits, and schema markup. Use when optimizing for AI search, generative search, or LLM visibility.

GitHub
安装命令
npx skhub add borghei/ai-seo
Markdown
SKILL.md

AI SEO

Generative engine optimization (GEO) for getting cited by AI search platforms — not just ranked in traditional results.


Table of Contents


Keywords

AI SEO, generative engine optimization, GEO, AI overviews, Google SGE, ChatGPT citations, Perplexity SEO, Claude citations, AI search optimization, semantic search, entity optimization, LLM visibility, AI-generated answers, structured data, schema markup, content extractability, AI citability, GPTBot, PerplexityBot, ClaudeBot, answer engine optimization


Clarify First

Before optimizing, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Target queries — the questions you want to be cited for (drives which pages to optimize and the extractable blocks to add)
  • Target AI platform(s) — Perplexity / ChatGPT / Google AI Overviews / Claude (crawling, indexing, and citation behavior differ per platform)
  • Brand/entity name — exact wording to track (drives citation testing and entity optimization)
  • The page/content to optimize — the URL or draft being restructured (drives extractability scoring + schema selection)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

Run an AI Visibility Audit

  1. Check robots.txt for AI bot access — the search/retrieval crawlers first (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot), then the training crawlers (GPTBot, ClaudeBot, Google-Extended)
  2. Test top 10 target queries on Perplexity, ChatGPT, and Google AI Overviews
  3. Document which queries cite you, which cite competitors, and what content format wins
  4. Score key pages against the Extractability Checklist
  5. Prioritize pages with highest gap between search volume and current AI citation presence

Optimize a Page for AI Citation

  1. Add a clear definition block in the first 200 words for informational queries
  2. Structure content with self-contained H2 sections that can be extracted independently
  3. Add numbered steps for process queries, comparison tables for "X vs Y" queries
  4. Replace all vague claims with attributed statistics ("According to [Source], [Year]")
  5. Implement schema that matches the page type (Article, Product, Organization, BreadcrumbList) — for accurate entity understanding, not as an AI-citation shortcut (Google states no special schema is needed for AI features)
  6. Verify AI bots are allowed in robots.txt

How AI Search Differs from Traditional SEO

The Fundamental Shift

Traditional SEO gets your page ranked. AI SEO gets your content cited. These are different optimization targets.

DimensionTraditional SEOAI SEO
GoalRank on page 1Get cited in AI-generated answers
Success metricClick-through rateCitation frequency
Content priorityKeyword densityAnswer extractability
Authority signalBacklinks + domain authorityBacklinks + answer quality + attribution
User interactionUser clicks your linkAI extracts your answer; user may never visit
Content formatLong-form comprehensiveSelf-contained extractable blocks
Optimization unitThe pageThe paragraph or section

What Carries Over from Traditional SEO

  • Domain authority still matters. AI systems prefer credible sources.
  • Backlinks still signal trust and expertise.
  • Technical SEO fundamentals (page speed, mobile-friendly, clean HTML) still apply.
  • Quality content with original insights still wins.

What Changes

  • Keyword density matters less than answer clarity and directness
  • Page-level optimization expands to section-level and paragraph-level optimization
  • Internal linking serves discoverability for AI crawlers, not just PageRank flow
  • Structured data helps machines disambiguate entities and page type, but it is not a gate: Google states there is no special schema.org markup required to appear in AI Overviews or AI Mode (as of September 2026)

The Three Pillars of AI Citability

Pillar 1: Structure (Extractable)

AI systems pull content in chunks. They find the paragraph, list, or definition that directly answers a query and extract it. Your content must be structured so answers are self-contained.

Extractability requirements:

  • Definition blocks for "what is X" queries — tight, 1-2 sentence definitions in the first 200 words
  • Numbered steps for "how to do X" queries — verb-first, self-contained steps
  • Comparison tables for "X vs Y" queries — clean table format with headers
  • FAQ blocks for question-based queries — explicit Q&A pairs
  • Statistics with full attribution for data-oriented queries

Anti-patterns that kill extractability:

  • Burying the answer in paragraph 8 of a 4,000-word essay
  • Requiring context from previous sections to understand any individual section
  • Using narrative prose for comparisons that should be tables
  • Placing key definitions only in the conclusion

Pillar 2: Authority (Citable)

AI systems do not just extract the most relevant answer — they extract the most credible one.

Authority signals in the AI era:

  • Domain authority — High-DA domains get preferential citation
  • Author attribution — Named authors with credentials outperform anonymous pages
  • Citation chains — Your content cites credible sources, making you credible in turn
  • Recency — AI systems prefer current information for time-sensitive queries
  • Original data — Proprietary research, surveys, and studies get cited more because AI cannot find this data elsewhere
  • Consistent entity presence — Your brand appears across authoritative sources as an entity

Pillar 3: Presence (Discoverable)

AI systems must be able to find and index your content.

Technical requirements:

  • AI crawlers allowed in robots.txt
  • Fast page load and clean HTML
  • No JavaScript-only rendering for important content
  • Schema markup for content type classification
  • Proper canonical signals
  • HTTPS with valid certificates

Core Workflows

Workflow 1: AI Visibility Audit

Step 1: Bot Access Verification

Check robots.txt for AI crawler permissions:

# Search / retrieval crawlers — blocking these removes you from that platform's AI answers:
Googlebot         # Google Search, incl. AI Overviews and AI Mode
OAI-SearchBot     # ChatGPT search
Claude-SearchBot  # Claude search
PerplexityBot     # Perplexity search
# Training crawlers — blocking these does NOT remove you from AI search answers:
GPTBot            # OpenAI model training
ClaudeBot         # Anthropic model training
Google-Extended   # Gemini training + grounding (not AI Overviews / AI Mode)
Applebot-Extended # Apple Intelligence training

If a search/retrieval crawler is blocked, that is the single highest priority fix — zero visibility on that platform until resolved. A blocked training crawler is a policy choice, not a visibility bug. See Bot Access Configuration for the full matrix.

Step 2: Citation Testing

Test top 10 target queries on each platform:

PlatformHow to TestWhat to Record
PerplexitySearch at perplexity.ai, check Sources panelCited? Which competitors cited? Content format winning?
ChatGPTWeb browsing enabled, check citationsSame
Google AI OverviewsGoogle query, check AI Overview panelSame
Microsoft CopilotSearch at copilot.microsoft.com, check source cardsSame
ClaudeWeb search enabled queriesSame

Step 3: Content Extractability Scoring

Score each key page (0-7):

  • Clear definition of core concept in first 200 words
  • Numbered lists or step-by-step sections for process queries
  • FAQ section with direct Q&A pairs
  • Statistics cited with source name and year
  • Comparisons in table format (not narrative)
  • H1 phrased as an answer or direct statement
  • Valid schema markup matching the page type (Article, Product, Organization, BreadcrumbList)

Interpretation: 0-3 = needs major restructuring. 4-5 = good baseline. 6-7 = strong.

Step 4: Competitive Citation Analysis

For each target query, document:

  • Who is currently being cited (top 3 sources per platform)
  • What content format wins (definition, list, table, quote)
  • What your content lacks that cited competitors provide
  • Where you have unique data or expertise competitors lack

Workflow 2: Page Optimization for AI Citation

Step 1: Lead with the Answer

The first paragraph must contain the core answer to the target query. No preamble, no context-setting, no "In today's landscape..." openers.

Step 2: Structure Self-Contained Sections

Every H2 section must be answerable as a standalone excerpt:

  • Each section opens with its main point
  • Each section contains its own evidence
  • No section requires reading previous sections to be understood
  • Each section could be quoted out of context and still make sense

Step 3: Add Extractable Content Blocks

Insert 2-3 of these per key page:

  • Definition block (first 200 words)
  • Numbered how-to steps (5-10 max, verb-first)
  • Comparison table (clean headers, structured data)
  • FAQ pairs (question matches natural language query)
  • Attributed statistics ("According to [Source] ([Year]), X% of...")
  • Expert quote block ("[Name], [Role at Organization]: '[quote]'")

Step 4: Replace Vague with Specific

Find and replace every vague claim:

  • "Many companies" → name the companies or cite the count
  • "Studies show" → name the study, organization, and year
  • "Significantly improved" → state the percentage improvement
  • "Leading brands" → name at least one
  • "Experts say" → name the expert with credentials

Step 5: Add Schema Markup

Implement JSON-LD in the page head:

Content TypeSchemaImpact (as of September 2026)
Articles and postsArticleMedium — Article rich result + clear author/date signals
Product pagesProduct (+ Review snippet)Medium — Product rich results; product comparison queries
Company pagesOrganizationMedium — entity authority, logo/knowledge panel signals
All pagesBreadcrumbListLow-Medium — Breadcrumb rich result, site hierarchy
Author pagesPerson / ProfilePageLow-Medium — author credibility signal
FAQ sectionsFAQPageLow — FAQ rich results removed from Google Search (May 2026); valid vocabulary, no Google search feature
Step-by-step guidesHowToLow — HowTo rich results removed from Google Search (2023); no Google search feature

The visible on-page structure (Q&A headings, numbered steps) is what AI systems extract. Google states no special schema is required for AI features — do not sell FAQPage/HowTo markup as an AI-citation lever.

Workflow 3: Entity Optimization

Step 1: Define Your Entity

Ensure your brand exists as a recognized entity across the web:

  • Wikipedia or Wikidata presence
  • Google Knowledge Panel
  • Consistent NAP (name, address, phone) across citations
  • Structured About page with Organization schema

Step 2: Build Entity Associations

Connect your entity to relevant topics:

  • Publish original research on topics you want to be cited for
  • Get mentioned (with links) on authoritative sites in your domain
  • Contribute expert quotes to industry publications
  • Maintain active presence on platforms AI systems index

Step 3: Strengthen the Citation Chain

Create a network of credible references:

  • Your content cites authoritative sources
  • Authoritative sources cite your content
  • Your author pages link to credentials and publications
  • Your brand appears in industry roundups and comparisons

Content Patterns That Get Cited

Pattern 1: Definition Block

**[Term]** is [concise definition in 1-2 sentences]. [One sentence of context
explaining why it matters or how it differs from related concepts].

Place within the first 200 words. No hedging, no preamble.

Pattern 2: Numbered Steps

Requirements for AI extraction:

  • Steps are numbered (not bulleted)
  • Each step starts with an action verb
  • Each step is self-contained (could be quoted alone)
  • 5-10 steps maximum (AI truncates longer lists)
  • Each step has a brief explanation (1-2 sentences)

Pattern 3: Comparison Table

Two-column or multi-column tables with clean headers:

| Dimension | Option A | Option B |
|-----------|----------|----------|
| Price | $X/mo | $Y/mo |
| Key Feature | Description | Description |
| Best For | Use case | Use case |

Pattern 4: FAQ Block

Explicit Q&A pairs. Questions should match natural language queries:

### What is [topic]?
[Direct answer in 1-2 sentences.]

### How does [topic] work?
[Step-by-step explanation.]

The visible Q&A structure is what matters for extraction. FAQPage markup is optional — Google no longer shows FAQ rich results (removed May 2026).

Pattern 5: Attributed Statistics

According to [Source Name] ([Year]), X% of [population] [finding].

Complete attribution is critical. Unattributed statistics get deprioritized because AI cannot verify the source.

Pattern 6: Expert Quote Block

"[Quote]" — [Name], [Role] at [Organization]

Named experts with credentials produce citable units AI systems pick up.


Schema Markup for AI Discovery

Priority Implementations

Google states that no special schema.org markup is needed to appear in AI Overviews or AI Mode (AI features and your website). Prioritize schema that still produces Google rich results or clarifies entities (Article, Product, Organization, BreadcrumbList); treat FAQPage and HowTo as optional vocabulary.

FAQPage Schema (optional — FAQ rich results removed from Google Search in May 2026):

{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "name": "What is [topic]?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "[Direct answer]"
      }
    }
  ]
}

HowTo Schema (optional — HowTo rich results removed from Google Search in 2023):

{
  "@context": "https://schema.org",
  "@type": "HowTo",
  "name": "How to [do thing]",
  "step": [
    {
      "@type": "HowToStep",
      "name": "Step name",
      "text": "Step description"
    }
  ]
}

Article Schema (still eligible for rich results; establishes author and date signals):

{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Title",
  "author": {
    "@type": "Person",
    "name": "Author Name",
    "url": "https://author-page"
  },
  "datePublished": "2026-01-15",
  "dateModified": "2026-03-01"
}

Validate all schema at schema.org/validator before deployment.


Bot Access Configuration

AI Crawler Matrix (as of September 2026)

The major AI vendors now split crawling by purpose, so search visibility and training can be controlled separately.

VendorUser agentPurposeBlocking it meansHonors robots.txt
OpenAIGPTBotModel trainingContent excluded from future trainingYes
OpenAIOAI-SearchBotChatGPT search indexNot surfaced/cited in ChatGPT search answersYes
OpenAIChatGPT-UserFetches a page when a user's request needs itUser-initiated; OpenAI says robots.txt rules "may not apply"Not guaranteed
AnthropicClaudeBotModel trainingFuture content excluded from trainingYes
AnthropicClaude-SearchBotClaude search indexReduced visibility in Claude search answersYes
AnthropicClaude-UserFetches a page when a Claude user asksClaude cannot retrieve your page for usersYes
PerplexityPerplexityBotPerplexity search indexNot surfaced/linked in Perplexity answersYes
PerplexityPerplexity-UserFetches a page for a user's questionPerplexity says this fetcher generally ignores robots.txtNo
GoogleGooglebotGoogle Search, including AI Overviews and AI ModeRemoves you from Google Search entirelyYes
GoogleGoogle-Extended (robots.txt token, not a separate crawler)Gemini training and grounding in some Google systems; also limits training of the models behind Search generative AI featuresDoes not remove you from Google Search, ranking, AI Overviews or AI ModeYes

Sources: OpenAI crawlers, Anthropic crawlers, Perplexity crawlers, Google common crawlers.

Google AI Overviews / AI Mode opt-out: Google-Extended does not control appearance in them — they are served from Googlebot's index. Your options (as of September 2026):

  • Search Console "Search generative AI" control (Settings > Search generative AI; rolled out to all sites August 31, 2026) — excludes the site's links and content from AI Overviews, AI Mode, and generative AI features in Discover, without affecting ranking or inclusion in the rest of Search. Takes effect within a few days (help).
  • Page-level snippet controls — nosnippet, data-nosnippet, max-snippet limit what can be shown, but also limit how the page appears in regular Search results.
  • noindex — removes the page from Google Search entirely.

Recommended robots.txt Configuration

# Allow AI search / retrieval crawlers (visibility)
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

# Training crawlers — Allow or Disallow per your content policy;
# this choice does not affect AI search citation
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

Training vs. Citation Access

Allowing AI citation while blocking training is enforceable for OpenAI, Anthropic, and Perplexity: allow the search crawler (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and disallow the training crawler (GPTBot, ClaudeBot). Blocking one does not block the other.

Google works differently: Google-Extended limits training and Gemini grounding but not appearance in AI Overviews or AI Mode. To keep regular Search visibility but leave AI Overviews / AI Mode, use the Search Console Search generative AI control (as of September 2026); snippet controls and noindex also work but cost regular Search presentation.

Recommendation: Always allow the search/retrieval crawlers if you want AI citation visibility; decide on training crawlers as a separate content-licensing policy.


Monitoring and Tracking

Weekly Citation Tracking (20 minutes/week)

Test top 10 target queries on Perplexity and ChatGPT:

  • Were you cited? (yes/no)
  • Citation rank (1st source, 2nd, 3rd)
  • What text was used from your content?
  • Any new competitors appearing?

Google Search Console for AI Overviews and AI Mode

There is no "AI Overviews" search-type filter in the Performance report. Clicks and impressions from AI Overviews and AI Mode are counted inside the standard Web search type, blended with regular results (AI features and your website).

As of September 2026, Search Console also has a dedicated Generative AI performance report (announced June 2026, rolled out to all sites by August 31, 2026; sites need enough impressions to see data) (announcement, help):

  • Impressions only — no click data — for links to your site shown in AI Overviews and AI Mode
  • Group by page, country, device, and date to see which pages get surfaced most
  • Standard Performance-report limits apply (1,000-row cap, preliminary recent data)

For clicks, use the Web search type in the Performance report plus landing-page analytics; AI-feature clicks cannot be isolated there.

Monthly Monitoring Checklist

SignalWhat to CheckTool
Perplexity citationsTop 10 queriesManual testing
ChatGPT citationsTop 10 queriesManual testing
Google AI Overviews / AI ModeImpressions (Generative AI report); clicks only blended into Web search typeGoogle Search Console
Copilot citationsTop 5 queriesManual testing
AI bot crawl activityCrawl frequency and pagesServer logs / Cloudflare
Competitor citationsWho is getting cited for your queriesManual testing
Content freshnessDate signals on key pagesContent audit

When Citations Drop

Diagnostic checklist when you lose a citation:

  1. Did robots.txt change? (Check for accidental AI bot blocks)
  2. Did a competitor publish more extractable content?
  3. Did your page structure change? (Restructuring can break citation patterns)
  4. Did your domain authority drop? (Check backlink profile)
  5. Did the query intent shift? (AI systems may reinterpret the query)

Best Practices

  1. Optimize at the section level, not just the page level — AI extracts paragraphs and sections, not entire pages. Every H2 block should be independently citable.

  2. Lead with the answer, always — The first 200 words determine whether AI systems find your content useful. Put the answer there.

  3. Attribute everything — Unattributed statistics, unnamed experts, and sourceless claims reduce your citability. Name names.

  4. Update quarterly — AI systems prefer recent content. Update publish dates and refresh data points every 90 days.

  5. Build entity presence — The stronger your brand's entity recognition across the web, the more AI systems trust and cite you.

  6. Do not choose between traditional SEO and AI SEO — They are complementary. Many optimization signals overlap. Run both.

  7. Test on multiple platforms — A page cited on Perplexity may not be cited on ChatGPT. Optimize for the platforms your audience uses.

  8. Monitor competitors monthly — Track who gets cited for your target queries and study what content patterns they use.

  9. Avoid JavaScript-rendered content for key answers — AI crawlers may not execute JavaScript. Ensure important content is in the initial HTML.

  10. Implement schema accurately, not as a hack — Article, Product, Organization, and BreadcrumbList markup that matches visible content helps entity understanding. Google says no special schema is required for AI features, and FAQ/HowTo rich results no longer display.


Integration Points

  • SEO Specialist — Use for traditional search ranking optimization. Run AI SEO and traditional SEO in parallel.
  • Content Production — Use to create the underlying content before optimizing for AI citation.
  • Content Humanizer — Use after writing. AI-sounding content performs worse in AI citations — AI systems prefer credible, human-sounding writing.
  • Content Strategy — Use when deciding which topics and queries to target for AI visibility.
  • Marketing Analytics — Use campaign analytics tools to track the business impact of AI citation traffic.

Troubleshooting

ProblemLikely CauseFix
Content not cited despite high DAPoor extractability — answers buried in proseRestructure with definition blocks, numbered steps, and FAQ pairs in first 200 words
Cited on Perplexity but not ChatGPTDifferent crawling and indexing pipelines per platformVerify bot access for all AI crawlers; test rendering without JavaScript
AI Overview shows competitor insteadCompetitor has more extractable, better-attributed contentAudit competitor's cited content format and match or exceed specificity
Citation dropped after site updatePage restructure broke the extraction pattern AI was usingCompare old vs new page structure; restore extractable blocks
OAI-SearchBot / PerplexityBot blocked in robots.txt unknowinglyCMS update or security plugin overwrote robots.txtAudit robots.txt after every CMS or plugin update; set up monitoring
Schema markup present but no rich resultsMissing required fields or content-markup mismatchValidate with Google Rich Results Test; ensure schema matches visible page content
AI cites your data but not your brandMissing entity signals — no Organization schema or sameAs linksImplement Organization schema with sameAs to Wikidata, LinkedIn, and social profiles

Success Criteria

  • AI citation rate: Achieve citation in 30%+ of target queries across Perplexity, ChatGPT, and Google AI Overviews within 90 days of optimization
  • Extractability score: Score 6-7 out of 7 on the Content Extractability Scoring checklist for all key pages
  • Bot access: Zero AI search/retrieval crawlers blocked in robots.txt (training crawlers per documented policy) — verified monthly with automated monitoring
  • Entity recognition: Brand appears in Google Knowledge Panel and is recognized as an entity on Wikidata
  • Schema coverage: 100% of content pages have JSON-LD schema matching the page type (e.g., Article, Product, Organization, BreadcrumbList) validated without errors
  • Freshness cadence: All key pages updated within the last 90 days with current dateModified signals
  • CTR from AI Overviews: Track organic CTR separately for queries where AI Overviews appear and hold it at or above your own pre-optimization baseline for those queries in Search Console

Scope & Limitations

In scope:

  • Optimizing content structure for AI extraction and citation
  • Bot access configuration and monitoring
  • Schema markup implementation for AI discoverability
  • Entity optimization and Knowledge Graph presence
  • Citation tracking across AI search platforms
  • Content pattern design (definitions, steps, tables, FAQs)

Out of scope:

  • Traditional organic ranking optimization (use SEO Specialist)
  • Content creation from scratch (use Content Production)
  • Paid search or paid AI placement strategies
  • AI model training data licensing or opt-out negotiations
  • Platform-specific API integrations for automated tracking
  • Social media optimization for AI-adjacent platforms

Known limitations:

  • AI citation tracking is largely manual — no standardized API exists across platforms
  • Citation algorithms are opaque and change frequently without notice
  • Blocking AI training while allowing citation works per vendor via separate user agents (OpenAI, Anthropic, Perplexity); Google AI Overviews / AI Mode use Googlebot, so appearance there is controlled in Search Console, not robots.txt
  • User-initiated fetchers (ChatGPT-User, Perplexity-User) may not honor robots.txt
  • AI summaries are reported to reduce clicks on traditional results (Pew Research Center, July 2025: users clicked a traditional result in 8% of visits with an AI summary vs. 15% without; Google reports overall organic click volume as relatively stable), and this cannot be fully mitigated

Scripts

# Analyze content for AI citability signals
python scripts/content_scorer.py page.html --json

# Simulate how content might appear in AI search results
python scripts/serp_simulator.py --query "what is cloud cost optimization" --content page.md

# Analyze keyword opportunities for AI search visibility
python scripts/keyword_analyzer.py --keywords keywords.csv --json
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

NOASSERTION

源路径

marketing/ai-seo

默认分支

main

最新提交

f308cbd

Tree SHA

d30ff9d