Most B2B teams starting an AEO programme make the same mistake. They ask ChatGPT or a prompt-research tool for the prompts to track. They accept the list and move on. It looks comprehensive. It covers the keywords. It sorts neatly into buckets. And it almost never moves the pipeline.
The reason is structural. AI-generated prompt lists optimize for category coverage, not buyer behaviour. They surface what a model thinks is relevant, not what your ICP actually types into ChatGPT, Perplexity, or Google AI Overviews during vendor evaluation. That gap is where most AEO programmes lose their ROI.
This article shows how to identify prompts that align with your ICP and convert. We use a 37-prompt OT security audit as the worked example and lay out the three-phase cycle that turns prompt selection into pipeline impact.
Why AI-Generated prompt lists fail B2B brands
What a typical AI tool gives you when you ask for prompts
Ask any LLM, "give me the top 30 prompts buyers to ask about OT security." The output will look something like this: What is OT security? Why is OT security important? What are the benefits of OT security? OT security vs IT security. Top OT security trends in 2026. Eight prompts.
All definitional. All awareness-stage. All vendor-agnostic. No persona, vertical, use-case, or comparison signal anywhere. This is the default output across cloud security, observability, MDM, dev tools, and every other B2B technical category.
Why those prompts will never drive qualified pipeline
The buyer asking "What is OT security?" is not a qualified prospect. That prompt comes from journalists, students, or junior analysts orienting before a meeting. Even if a brand surfaces in that answer, nothing converts. There is no purchase intent in the prompt.
The prompts that convert sound completely different. A job title ("I'm a CISO"), an industry ("for pharmaceutical manufacturing"), a use case ("rugged firewall for harsh environments"), a competitor name ("Fortinet vs Cisco"), or a compliance trigger ("IEC 62443 implementation").
Forrester's 2026 Buyers' Journey Survey of 18,000 global buyers found that generative AI now outranks vendor websites and sales reps as the most meaningful purchase research source. If the tracked prompts carry no buyer signal, the brand is invisible at the exact moment buyers are building their shortlist inside AI conversations.
Without persona, vertical, or comparison signal in the prompt, the tracking measures awareness for the sake of awareness.
How do AI tools generate Prompts?
AI prompt generators work by pattern-matching against training data. The model scans for the most statistically common questions associated with a topic and outputs a frequency-ranked list. The result reflects what appears most often across the internet, not what a specific buyer persona asks during vendor evaluation.
This matters because training data skews heavily toward awareness-stage content. Blog posts, Wikipedia entries, and beginner guides dominate the web for most B2B categories. Decision-stage content like vendor comparisons, compliance implementation queries, and role-specific evaluations is far less common. The model generates what it has seen the most of. And that is top-of-funnel.
The same pattern holds for AEO platforms with auto-suggest features. Most pull from keyword databases built for traditional search, not from buyer conversations. A prompt with 10,000 monthly searches and zero conversion potential gets included. A prompt with 50 searches and strong CISO buying signal gets excluded.
No AI tool has access to sales discovery calls, win/loss interviews, or the actual language buyers use when shortlisting vendors. That language lives inside CRM notes, call recordings, and Reddit threads. Until those inputs are layered in, any auto-generated list reflects category popularity rather than ICP behaviour.
Generic Prompts vs Prompts That Convert
Here is a direct comparison from the OT security category. The left column is what an AI tool produces with a generic "give me OT security prompts" request. The right column is what we tracked in the actual 37-prompt audit.
The pattern is consistent. Generic prompts produce educational answers. ICP-aligned prompts produce vendor shortlists. The first fills a dashboard. The second fills a pipeline.
5 Dimensions for Comprehensive, Data-Backed Prompt Coverage
Every prompt that drives meaningful AEO outcomes carries signal in at least one of five dimensions. A complete prompt set covers all five. Each dimension corresponds to a different moment in the buyer's research journey.
1. Persona-specific The buyer's role appears in the prompt. The AI returns role-relevant recommendations instead of category-level answers. "As an OT security engineer, what tools should I use for ICS threat detection?", "I'm a CISO. How should I evaluate OT security platforms?", "As an IT head, what should I look for when choosing an OT security vendor?"
2. Industry and vertical-specific The vertical appears in the prompt. The AI narrows its recommendations to vendors serving that sector. "Best OT security solutions for manufacturing", "OT platforms for oil and gas", "Which OT security solutions work best for pharmaceutical manufacturing compliance?"
3. Use-case and product-fit The buyer is solving a specific problem with specific constraints. The AI responds with product-level answers, not category overviews. "What's the best rugged firewall for harsh OT environments?", "Which SIEM is best for monitoring OT and ICS networks?", "How does NAC help with OT device visibility?"
4. Compliance and framework A regulation or standard is the trigger. Compliance has a deadline, so intent is high. "What is IEC 62443 and how do I implement it?", "What does the NIST cybersecurity framework recommend for OT?", "How does zero trust apply to OT environments?"
5. Comparison and competitor A specific competitor is named. The AI picks a side or ranks vendors head to head. These are decision-stage prompts and the highest-leverage ones in any tracking set. "How does Fortinet compare to Claroty?", "Fortinet vs Cisco for OT security", "Fortinet vs Palo Alto Networks for OT"
Use Case: Building a 37-Prompt Set for a Cybersecurity Company
The OT security audit illustrates the methodology end to end. The final prompt set did not come from a category brainstorm. It came from a two-week sourcing process across four inputs:
- Sales discovery calls produced persona-bound prompts. "I'm a CISO" and "as an OT security engineer" were verbatim from buyer conversations.
- Google Search Console, filtered for non-zero impressions and below-average CTR, produced the compliance and framework prompts.
- Ahrefs question and comparison keywords produced the head-to-head competitor prompts.
- Reddit threads across r/OTSecurity, r/sysadmin, and r/networking produced use-case prompts around legacy devices, asset visibility, and rugged firewalls.
No single source would have produced the full set. An AI tool would have produced none of it.
The 37 prompts were distributed across five dimensions:
Three prompts illustrate why sourcing discipline matters:
- "How do I secure legacy OT devices that can't be patched?" came from Reddit. An AI tool would have returned "OT patch management."
- "What's the best OT security platform that covers firewalls, NAC, and SIEM together?" came from a sales call. An AI tool would have returned "OT security platforms."
- "Which OT security solutions work best for pharmaceutical manufacturing compliance?" combines vertical, compliance, and product intent in one prompt. An AI tool splits that into three generic queries and loses the signal entirely.
The full OT security prompt tracking template covers all 37 prompts with sourcing, category, and visibility outcomes mapped across each one.
The three-phase ICP-aligned prompt cycle
An ICP-aligned prompt programme is not a one-step exercise. The work runs across three connected phases, each with its own inputs, decisions, and outputs. Skipping any one of them produces dashboards that look complete but cannot be acted on.
Phase 1 β Identify Prompts Your Buyers Actually Use
The first phase produces the ICP-aligned prompt set. The five-dimension framework (persona, industry, use case, compliance, comparison) is the lens. The inputs come from four sources:
- Sales call mining: Discovery calls produce persona-bound prompts in literal buyer language. Phrases like "I'm a CISO" and "as an OT security engineer" come verbatim from these calls.
- GSC query alignment: Queries with non-zero impressions and below-average CTR are buyer questions LLMs are now answering instead of the site.
- Ahrefs search-volume validation: Volume confirms whether buyers are actually searching the topic. A prompt with zero volume and no sales-call match is hypothetical and should be dropped.
- Reddit and community sourcing: Subreddits and industry forums surface use-case and pain-point prompts in raw practitioner language that no AI tool would generate.
The output is a prompt list validated against three things. Real buyer language. Real search behaviour. And real conversion goals that map back to revenue targets, not generic category coverage. A prompt with both search volume and a verbatim sales-call match is the strongest signal in any tracking set.
Phase 2 β Map Visibility Across Every Prompt
Once the prompt set is in place, Phase 2 runs the analysis that tells the team what to do next. Three outputs come out of this phase:
- Per-prompt visibility: Is the brand mentioned, cited with a link, both, or neither?
- Competitive landscape: Which competitors are winning, which prompts, and on which AI platforms?
- Dominant content mode: For every cited answer, where does the citation live? On the vendor's owned domain, on an earned source (industry publication, third-party listicle, peer review), or on a social platform (Reddit, LinkedIn, YouTube)?
The third output is the one most teams skip. It determines the entire content response plan. If the dominant mode for a prompt is earned media, building another owned page will not move the needle. If the dominant mode is social, a corporate blog post will not surface in the answer.
This is not theoretical. Muck Rack's analysis of 25 million AI citations found that earned media accounts for 84% of all AI citations. Paid and advertorial content accounts for just 0.3%. Without citation-source mapping in Phase 2, most teams default to "write more owned content" and miss where the visibility actually lives.
Phase 3 β Turn Gaps Into Content That Wins
Phase 3 turns the analysis into a prioritised action plan. Every prompt gets one of four responses, dictated by the dominant mode identified in Phase 2:
- Owned media dominant, brand already cited: Optimise. Tighten the answer-first lead, add a comparison table, refresh the data.
- Owned media dominant, brand not cited: Build. A new glossary entry, a vertical solution page, a comparison battle page.
- Earned media dominant: Owned content alone will not move the needle. The play is direct outreach to listicle authors, contributed columns in industry publications, or G2 and Gartner Peer Review contributions.
- Social platforms dominant: The play is distribution. A long-form LinkedIn article from a recognised practitioner, a Reddit AMA in the relevant subreddit, a YouTube technical walkthrough.
The OT security audit illustrates this clearly. Vertical-specific prompts where the brand was already winning needed optimisation, not new content. CISO-evaluation prompts where competitors dominated needed earned plays. Use-case prompts where Reddit was being cited needed practitioner-voice content, not corporate marketing pages. Three different prompts, three different content responses. All driven by Phase 2 citation mapping.
Compounding AI-Search Visibility Every 90 Days
Phase 3 is not the end. Buyer language shifts. Competitors publish new content. LLMs re-rank citation sources every quarter. The framework is built as a cycle for that reason.
The prompts validated in Phase 1 should be re-tested in Phase 2 every 90 days. Earlier if a major competitor moves on a comparison prompt. The action plan from Phase 3 informs the next round of Phase 1 sourcing. Prompts that did not move despite content investment go back to the drawing board. Prompts that started winning produce templates to apply elsewhere.
This is the structural reason ICP-aligned programmes outperform AI-generated lists. AI-generated lists are static. They do not feed forward. The three-phase cycle is iterative, and every quarter sharpens the dashboard further.
The Prompt Methodology: From Broad Category to High-Intent Coverage
Step 1: Mine sales calls for the literal phrasing. Pull 30 to 50 recent discovery calls. Read the first ten minutes of each. Note the actual phrasing buyers use when describing their problem and their role. Most persona-specific and use-case prompts come from this step. The most common mistake is teams paraphrasing buyer language into clean marketing copy. Do not. Buyers type the messy version into ChatGPT, not the polished one.
Step 2: Validate with GSC and Ahrefs.Take the candidate prompts from Step 1 and check them against GSC impressions and Ahrefs keyword volume. Prompts with non-zero impressions and below-average CTR are buyer questions LLMs are now answering instead of the site. Prompts with zero traffic anywhere are hypothetical. Drop them unless a sales call confirms the language.
Step 3: Tag every prompt before it enters the tracking set. Map each prompt across four axes: persona (CISO, IT head, engineer, plant manager), vertical (manufacturing, energy, pharma), use case, and funnel stage. The tagging takes 30 minutes for 50 prompts. It pays off every time the dashboard is sliced later.
Step 4: Verify coverage against conversion goals. Before signing off on the prompt set, run it past one question. If every prompt on this list converted to a strong win, would the resulting traffic match pipeline targets? If the prompts skew awareness, the answer is no. That is a coverage problem to fix before any tracking begins. This is the step most teams skip and the reason most AEO dashboards drift toward vanity metrics.
How Many Prompts Are Enough? Validating Your Prompt Set
Before any tracking goes live, the prompt set should cover six things:
- All named personas in the ICP.
- Verticals that represent more than 10% of revenue.
- Products and use cases the company actively sells.
- Named competitors from win/loss reports.
- Compliance frameworks relevant to the category.
- At least three head-to-head comparison prompts.
Missing any one of these creates a gap. Gaps in tracking become gaps in the dashboard, which become gaps in pipeline attribution.
Three gaps repeat almost every time a self-built prompt list gets audited:
- Missing personas: Most lists track products, not buyers. The CISO and IT-head perspectives are absent entirely.
- Missing verticals: Lists tend to over-index on the largest vertical and ignore the next two.
- Comparison prompts capped at one or two: Comparison is decision-stage. One or two prompts is not enough coverage to spot a competitor-favoured third-party article before it starts pulling pipeline.
Right-Fit Buyer Prompts: A Marketing Leader's New Mandate
How to brief an AEO programme owner. Stop asking the AEO platform or content team for "the top 50 prompts in our category." The output will be AI-generated and ICP-blind. Brief on coverage instead:
- All personas in the ICP.
- All verticals above 10% of revenue.
- All product lines.
- All named competitors.
- All relevant compliance frameworks.
Put a target on each dimension. The brief should look like a sourcing spec, not a wishlist.
Tying prompt-set quality back to pipeline metrics. The metric that matters is pipeline created from AI-influenced traffic, segmented by prompt dimension. When the dashboard shows that vertical-specific prompts convert at three times the rate of definitional prompts (and they will), the budget split for content investment writes itself. An ICP-aligned prompt set is not just better tracking. It is a sharper signal back to the rest of the marketing org about where the pipeline is hiding.
Future-Proofing Your Prompt Set: Keeping Pace With AI Search
An ICP-aligned prompt set compounds because it is built from real buyer behaviour. Buyer behaviour does not change overnight. The OT security set built in week one was 80% stable a quarter later. Only the comparison and emerging-trend prompts needed a refresh. AI-generated prompt sets do the opposite. They drift fastest because they were never anchored to anything real in the first place.
The shortcut of asking an LLM for a prompt list is tempting. It saves two weeks of sourcing. It produces a clean dashboard immediately. And it almost never moves the pipeline.
The two-week investment in persona, vertical, use-case, compliance, and comparison sourcing, combined with the three-phase Identify-Analyse-Act cycle, is what separates an AEO programme that drives leads from one that decorates a quarterly slide. This is the framework LeadWalnut runs for enterprise clients, rated 4.9 on Clutch across verified reviews.
FAQ
Can ChatGPT generate the initial prompt list for refinement?
As a starting point only. Never as the final list. AI-generated lists almost always skew awareness-stage and miss persona, vertical, and comparison signal. The AI output should account for 10% of sourcing, not 100%. The other 90% comes from sales calls, GSC, Ahrefs, and Reddit.
How many prompts should an ICP-aligned set contain?
Between 50 and 150 prompts per category, distributed across the five dimensions. The OT security worked example used 37 prompts because it was scoped tightly to a single sub-category. A broader portfolio with multiple product lines and verticals will need 80 to 150.
How long does the three-phase cycle take to run end to end?
Phase 1 (Identify) takes two weeks for a fresh category. Phase 2 (Analyse) takes 3 to 5 days once the prompt set is fixed. Phase 3 (Act) is content-production-bound, typically 4 to 8 weeks for a meaningful intervention. Re-audits at 90 days are lighter. The prompt set is already in place, so the work is mostly Phase 2 and Phase 3 updates.
How often should the prompt set be refreshed?
Re-audit the existing set every 90 days. Refresh the set itself every 180 days. Most prompts hold up across quarters. Comparison and emerging-trend prompts shift fastest.
How is this different from traditional keyword research?
Traditional keyword research optimises for search volume on Google. ICP-aligned prompt research optimises for the prompts buyers type into LLMs during evaluation. GSC search volume is one input, but the more important input is the literal phrasing buyers use. That phrasing often has zero traditional search volume but real LLM activity behind it.

How can LeadWalnut help?
Related Articles

Top B2B Content Creation Agencies Built for SaaS Pipeline in 2026

Perplexity SEO: How B2B Brands Earn Citations In AI Search



