Structured Data, Business Profile & Listings: The Technical Base for AI Citation
This is the one technical piece in the cluster. The others are the "why" and the "what" — this is the "how to set it up": schema.org, NAP consistency, Business Profile + Bing Places, and the new layer (llms.txt, robots OAI-SearchBot). For webmasters, developers, agencies — and the curious owner.
Why AI "digests" a structured page better
When an AI engine — ChatGPT, Gemini, Perplexity, Copilot — decides which venue to recommend, it doesn't look at your page design. It tries to read the facts out of the raw HTML: where you are, when you're open, how many stars you have, what's on the menu. The less it has to "guess," the more likely it is to cite you. That's exactly what structured data (schema.org / JSON-LD) does: it spells out the facts in machine language, unambiguously.
A widely-quoted figure floats around the market — schema-marked pages get cited roughly 3.2x more often, plus a ~73% selection advantage over unmarked ones. Handle it with care: this traces back to a BrightEdge-attributed claim with no verifiable primary report — treat it as an unverified vendor number, not a law (the closest verifiable BrightEdge figure is a 44% citation increase for structured data plus FAQ). The honest, verifiable picture is more nuanced.
In a causal study (1,885 pages adding JSON-LD vs. 4,000 controls, Aug 2025–Mar 2026), Ahrefs found no meaningful uplift on already-cited pages: AI Mode +2.4%, ChatGPT +2.2% (noise), AI Overviews −4.6%. And Growth Marshal (730 AI citations, 1,006 pages, Feb 2026) showed that generic schema alone doesn't help (in fact, a 41.6% citation rate, below the 59.8% for no schema) — only attribute-rich schema (real prices, hours, AggregateRating) lifted the citation rate to 61.7%, especially for low-authority domains. The lesson isn't that schema is useless; it's that the content must fill the schema with real, verifiable facts.
The core hospitality schemas, in order: LocalBusiness schema for a restaurant
These are the schema.org types worth putting on your site as JSON-LD. Crucial: AI crawlers parse JSON-LD from the raw HTML without executing JavaScript — so it must be server-rendered, not injected later with JS.
- →LocalBusiness / Restaurant — the foundation. Name, address (PostalAddress), phone, GeoCoordinates, price range, opening hours (OpeningHoursSpecification). Use `Restaurant` for a restaurant, `CafeOrCoffeeShop` for a café, `BarOrPub` for a bar.
- →Menu / MenuItem — the menu in machine form, with dietary flags (vegetarian, gluten-free). Your GBP menu is also a direct AI input, but the on-site Menu schema is your own controlled source.
- →AggregateRating + Review — the overall star score and review count. IMPORTANT: populate it with REAL ratings, never invented data — fake review schema is a policy violation and backfires. This field is fed precisely by your fresh Google reviews.
- →FAQPage — common questions in structured form ("Do you have a terrace?", "Do I need to book?", "Are you dog-friendly?"). These map directly to the question-based queries AI engines run.
- →OpeningHoursSpecification + GeoCoordinates + BreadcrumbList — hours, exact coordinates and the navigation breadcrumb. AI looks for hours and address as verifiable facts during selection.
Across the whole web, only ~12.4% of domains use any schema at all — making this a surprisingly easy edge that most of your competitors haven't picked up yet. (As of January 2026 Google is deprecating some legacy schema types, so build on the evergreen ones listed above.)
NAP consistency is the silent knock-out factor
NAP = Name, Address, Phone. If these don't MATCH across every surface (website, Google Business Profile, Bing Places, Facebook, TripAdvisor), AI would rather stay silent than state wrong data — and drops you from the running. One analysis found ~80% of local AI recommendations come from businesses with synchronized data, while ChatGPT/Perplexity data accuracy is only ~68% (Gemini is 100%, because it's grounded directly in Google Maps). A single misspelled street name or stale phone number is enough to keep you invisible.
Filling out Google Business Profile — through an AI lens
Your Business Profile is the foundation of all local visibility: Whitespark 2026 puts GBP signals at ~32% of Local Pack ranking weight (review signals at ~20%), and Gemini is essentially 100% grounded in Google's data. Complete all of these:
- ✓Exact primary category — not just "Restaurant" or "Bar," but the most precise subtype (Cocktail bar, Wine bar, Café, Pub, Night club). ~86% of profile views come from category- and attribute-based searches; the wrong category = exclusion from filtered queries.
- ✓Every relevant attribute — terrace, live music, happy hour, late-night, dog-friendly, good for groups, takes reservations. An unchecked attribute = automatic exclusion from AI questions like "dog-friendly café with a terrace."
- ✓Fresh photos + correct hours — the Gemini-powered "Ask Maps" (March 12, 2026) builds its recommendation in real time from the profile, reviews and website; current photos and hours improve your odds.
- ✓Reviews and responses — Google's own Business Profile help confirms that more reviews and positive ratings come with higher ranking, and that responses help you stand out. But Google officially frames replying as a customer-service gesture, not a guaranteed ranking lever — present it as a freshness and trust signal, not a promised algorithmic boost.
Bing Places: don't skip it if ChatGPT/Copilot is the goal
Here's a counterintuitive but important nuance. ChatGPT runs on the Microsoft Bing index, not directly on Google — and research suggests ChatGPT does not directly read Bing Places business profile fields; it reads the Bing-indexed open web. Seer Interactive found that 87%+ of SearchGPT citations matched Bing's top-20 organic results (vs. only 56% for Google).
So filling out Bing Places doesn't by itself "put you in ChatGPT," but it's still not skippable for two reasons: first, Microsoft Copilot's primary local source is the Bing index + Bing Places (relaunched Oct 3, 2025, with an AI recommendation tool for missing hours, photos and categories). Second, a complete Bing profile strengthens the data consistency of the Bing ecosystem that indexation is embedded in.
Good news that takes the pressure off: Bing rank is a weak predictor — Bing's top-3 URLs matched actual ChatGPT citations only 6.8–7.8% of the time. You don't need to "win" Bing's rankings. You just need to be indexed, with structured, verifiable, review-rich content. That sets the bar far lower than you'd think.
The new layer: llms.txt and robots.txt AI rules
In 2026 webmasters keep asking: do I need an llms.txt? Should I block AI bots in robots.txt? The short, honest answer: don't count on llms.txt as a GEO tool, and configure robots.txt deliberately — by crawler job.
Myths / what NOT to do
- Believing llms.txt gets you into AI recommendations — no major engine reads it in production (1 of 94,614 cited URLs pointed to an /llms.txt page). Google (Mueller, Illyes) explicitly does NOT support it, comparing it to the keywords meta tag.
- Blindly blocking "AI bots" with a single rule — you can lock out the search/citation bots too.
- Fearing a Google-Extended block hurts your Google ranking — AI Overviews run on Googlebot's index, not Google-Extended, so blocking it does not hurt your Google ranking.
- Assuming robots.txt definitely blocks — it's voluntary; on Dec 9, 2025 OpenAI removed ChatGPT-User from robots.txt compliance.
What to ACTUALLY do
- ALLOW the citation/search bots: OAI-SearchBot, ChatGPT-User, PerplexityBot, Googlebot, Bingbot — so they can cite you.
- Optionally BLOCK training-only bots (GPTBot, Google-Extended, CCBot, Bytespider) — purely an IP/training stance, it does NOT reduce AI-search visibility.
- Cloudflare user? Check your dashboard: since July 1, 2025 new domains block AI crawlers by default — you may be invisible without knowing it.
- If you do publish an llms.txt (cheap, harmless future-proofing), set it to noindex — but know it's optional and low-priority at best today.
Crawlable HTML and speed: if it can't read you, it won't cite you
Schema only matters if the bot can reach and read it. AI crawlers parse JSON-LD from the raw HTML — if your page content only appears via client-side JavaScript, the bot may see an empty page. Server-side rendering (or static HTML) is the safe choice.
Freshness is a technical signal too: since GPT-5 (Aug 2025) OpenAI tripled (3.5x) its real-time web crawling (the OAI-SearchBot surge), and on Dec 9, 2025 removed the word "training" from its SearchBot docs — meaning ChatGPT now leans far more on the live web than on static training data. According to SOCi, content updated within 30 days gets 3.2x more citations. A fresh, dated, well-structured page is simply a competitive edge.
Don't forget query fan-out: ChatGPT runs an average of 2.1 sub-queries per prompt and silently injects words like "best," "reviews" and the current year. A page covering multiple angles (menu + ratings + hours + neighborhood) outranks a page optimized for a single phrase — another reason well-structured, rich content wins.
A copy-paste technical checklist + where it fits in the GEO stack
If, as a webmaster or agency, you run this list through a venue, the technical base is solid. But keep the order straight: schema is the icing — the cake is reviews and a complete Business Profile.
- ✓JSON-LD server-side — LocalBusiness/Restaurant + Menu + FAQPage + AggregateRating/Review, populated with real data.
- ✓NAP match everywhere — website, Google Business Profile, Bing Places, Facebook, TripAdvisor; one canonical name/address/phone.
- ✓Google Business Profile 100% — exact category, every attribute, fresh photos, correct hours, active responses.
- ✓Bing Places filled out — mainly for Copilot; no need to win rankings, just be consistent.
- ✓robots.txt by crawler job — search/citation bots allowed; Cloudflare setting checked.
- ✓Crawlable, fast HTML — server-rendered, fresh and dated content.
- ✓The real fuel: reviews — the schema's AggregateRating/Review fields and the Business Profile signals are filled by fresh, plentiful, actively-answered Google reviews. This is exactly what ThanksBot automates: collection (e.g. via QR codes) and responding in 50+ languages to every review. It doesn't "put you in ChatGPT" — it strengthens the public review signals that AI engines read.
For the deeper background — why these signals matter — see the GEO and review pieces: Google Business Profile and local SEO (/blog/google-cegprofil-lokalis-seo-elso-hely), QR-code review collection (/blog/qr-kodos-velemenygyujtes-google-ertekeles) and how Google reviews affect restaurant revenue (/blog/google-velemenyek-hatasa-ettermi-bevetel-2025).
Frequently Asked Questions
What is LocalBusiness/Restaurant schema, and do I need it on my restaurant's website?
LocalBusiness (and its subtypes Restaurant, CafeOrCoffeeShop, BarOrPub) is a schema.org standard that lets you state the facts in machine-readable, unambiguous form (JSON-LD): name, address, phone, hours, price range, coordinates. Yes, it's worth adding: AI crawlers read it from the raw HTML, so you reduce ambiguity and raise your chances of being cited. It's not a magic bullet on its own — you must fill it with real, verifiable data — but since barely ~12.4% of domains use any schema, it's an easy edge.
How does structured data help AI cite me?
During selection, AI engines look for verifiable facts (hours, address, stars, attributes). Schema serves these up in machine form so the engine doesn't have to guess. Important nuance: causal studies (Ahrefs) show generic schema gives no meaningful uplift on already-cited pages — but attribute-rich schema (real prices, AggregateRating) does, especially for lower-authority domains. So schema helps your already-good content be read accurately and confidently by AI.
What is NAP consistency and why does it matter?
NAP is Name–Address–Phone. It's consistent when it appears exactly the same on every surface: website, Google Business Profile, Bing Places, Facebook, TripAdvisor. If it differs, AI would rather stay silent than state wrong data — effectively a knock-out factor. An estimated ~80% of local AI recommendations come from businesses with synchronized data, and ChatGPT/Perplexity data accuracy is only ~68% anyway, so every discrepancy hurts your odds.
Do I need a Bing Places account alongside Google Business Profile?
If ChatGPT or Microsoft Copilot is your goal, yes — with nuance. ChatGPT runs on the Bing index, and Copilot's primary local source is the Bing index + Bing Places (with an AI recommendation tool since Oct 2025). Filling out Bing Places doesn't "put you in ChatGPT," and Bing rank is a weak predictor (its top-3 matches actual citations only ~6.8–7.8% of the time), but consistent, indexed presence is a prerequisite. Google Business Profile remains the foundation (Gemini is essentially 100% grounded in it).
What is llms.txt, and does my restaurant website need one?
llms.txt is a proposed file meant to "show" AI models your important content. In 2026, though, no major AI engine reads it in production: Google (Mueller, Illyes) explicitly doesn't support it, and server logs show crawlers almost never request it (1 of 94,614 cited URLs pointed to one). For a restaurant/café/bar it's low priority. If you publish one (cheap, harmless future-proofing), set it to noindex. The real technical levers are a targeted robots.txt and complete profiles.
Can I set up schema without coding?
Partly, yes. Many website builders (WordPress plugins, Wix, Squarespace) and local-SEO tools can generate LocalBusiness/Restaurant JSON-LD with no code. The key is that the schema lands in the raw, server-side HTML (not just injected via JS) and contains REAL data — especially attribute-rich (prices, hours, AggregateRating). And the review side (the source of AggregateRating/Review) is brought to life by fresh, plentiful, answered Google reviews — which ThanksBot automates, no coding required.
Related Articles
Leave Your Google Replies to ThanksBot
14-day free trial. No credit card required. Cancel anytime.