Does ChatGPT Make Up Hotels? ChatGPT vs Gemini vs Perplexity Rates 2026
Trip Planning·11 min read·August 27, 2026

Does ChatGPT Make Up Hotels? ChatGPT vs Gemini vs Perplexity Rates 2026

Does ChatGPT Make Up Hotels? ChatGPT vs Gemini vs Perplexity Rates 2026

By Rachel Caldwell, AI Travel Editor at Travel Anywhere. Editorial verification August 25, 2026.

Last updated: 2026-08-25

Yes, it can, and the ways it goes wrong are more varied than most travelers expect. You ask for a boutique hotel in Lisbon and get a name, a neighborhood, a price band and a description that reads perfectly, and then it returns nothing on Booking.com, nothing on Google Maps, and nothing anywhere else. That is the dramatic version. The quieter versions cost more often: the rooftop pool that does not exist, the Marriott Courtyard placed in the wrong city, the property that closed in 2023 and still gets recommended in 2026, the amenity list assembled from what a hotel of that type usually has.

Four pains, four answers in this guide. The invented property is answered by the question of whether ChatGPT makes up hotels and how often. The rooftop pool that is not there is answered by the mechanism section on why models invent hotels. The Courtyard in the wrong city is answered by how real travelers catch a fake hotel before booking. The hotel that closed three years ago is what the Travel Anywhere verification stack catches.

This is not an edge case in the abstract. In January 2026, an AI-written blog post on the website of the Tasmanian operator Tasmania Tours recommended "Weldborough Hot Springs" in northeast Tasmania as a forest escape. There are no hot springs at Weldborough. Kristy Probert, who owns the Weldborough Hotel, told reporters she began getting five phone calls a day and two to three in-person visitors asking for directions to them. The company's owner, Scott Hennessey, said flatly: "our AI has messed up completely."

TL;DR: No clean per-model "hotel hallucination percentage" exists in published research, and anyone quoting one without a source is inventing it. What does exist: the Tow Center for Digital Journalism's March 2025 study of eight AI search products found Perplexity had the lowest citation error rate at 37%, ChatGPT Search 67%, and Grok 3 94% [SYNTHESIS], on news attribution rather than hotels. Vectara's leaderboard scores base models on grounded document summarization, which is a different task again. Real fabrication cases are documented by CNN, ABC News Australia, Fodor's and a January 2026 Journal of Consumer Behaviour paper. The decision rule: treat every AI hotel name as a lead, not a booking, until two independent platforms confirm it.

Editor's verification, Travel Anywhere desk: our editors re-checked this post's statistics against their primary sources on August 25, 2026. We confirmed the 37% / 67% / 94% figures, the authors and the 1,600-query methodology directly in the Tow Center's own published article rather than through a secondary summary. We read Vectara's leaderboard on GitHub and confirmed its stated scope. We confirmed the names and quotations in the Weldborough case against ABC News Australia's and CNN's reporting. Two claims that appeared in the earlier draft could not be confirmed and have been removed: a three-minute verification timing attributed to hospitality security research, which was our own editorial estimate, and a 76.7% financial-domain hallucination figure whose source we could not locate.

Key Takeaways

  • No peer-reviewed study has published a hotel-specific hallucination rate for ChatGPT, Gemini or Perplexity. That absence is the single most important fact on this page, because every precise per-model hotel percentage in circulation is extrapolated from benchmarks that tested something else [SYNTHESIS]. (source: Vectara hallucination leaderboard, read August 25, 2026)
  • Perplexity had the lowest citation error rate of eight AI search tools tested, at 37%, against ChatGPT Search at 67% and Grok 3 at 94%, in the Tow Center for Digital Journalism's March 2025 study of 1,600 news-attribution queries [SYNTHESIS]. That is news citation, not hotel recommendation, and 37% still means more than one attribution in three goes wrong. (source: Columbia Journalism Review / Tow Center, March 2025)
  • Documented fabrication cases exist across platforms and are now on the record with names attached. The Weldborough Hot Springs case was reported by ABC News Australia and CNN in January 2026, with the hotel owner and the tour company owner both quoted. (source: CNN Travel, January 2026)
  • Peer-reviewed tourism research names the failure types. A January 2026 Journal of Consumer Behaviour paper identifies "fictitious attraction opening hours or references to non-existent restaurants" as characteristic GenAI travel hallucinations, and measures the damage they do to perceived accuracy and trust. (source: Journal of Consumer Behaviour, January 2026)
  • Risk is highest for boutique, independent and recently opened properties. Thin web presence means less signal for the model to work from, and details like seasonal hours and ownership changes go stale fastest at small properties. (source: eHotelier, December 2025)
  • Three checks on two booking platforms plus Google Maps will settle almost every case. If a property does not appear on all three with matching details, do not build a trip around it. (Editorial method, Travel Anywhere desk, August 2026 [PLATFORM])

Read the full AI travel hallucination guide

Does ChatGPT Make Up Hotels, and How Often?

Yes, it can, and nobody knows how often. ChatGPT generates hotel names and descriptions from learned patterns and has no internal mechanism to confirm a property it names actually exists. The risk is highest for boutique, independent and obscure properties with thin web presence; major branded hotels in well-indexed cities are rarely invented outright, but wrong details attached to real properties are common even there. No published study reports a hotel-specific hallucination rate for ChatGPT or any competitor.

That second half is not a hedge, it is the finding. Here is what the available evidence does and does not support:

Model What the published evidence actually says Hotel error types documented in reporting How its errors fail Verification priority
ChatGPT with search enabled 67% citation error rate on news attribution in the Tow Center's March 2025 test [SYNTHESIS]. No hotel-specific figure published Invented boutique properties with plausible names; wrong amenities on real hotels; outdated pricing; closed properties listed as open Fluently and without flags High: cross-check every name
ChatGPT without browsing Not separately measured. Generates from training data with no live source and no recency Invented boutique properties, wrong amenities on real hotels, outdated pricing, plus staleness with no correction path Silently, with no citation to check Highest: do not use for hotels
Gemini No published citation-accuracy figure for hotel queries. Google integration gives it live data access it does not always use Wrong star rating or category; fabricated amenities; incorrect neighborhood mapping; the Weldborough case involved an AI content tool on a business website Confidently, often with a real-looking source High, especially outside major chains
Perplexity Lowest citation error rate of eight tools at 37% in the same Tow Center test [SYNTHESIS] Cites a real live URL and attaches a claim that page does not make; fewer invented names, more attribution errors Auditably, if you open the link Medium-high: read the cited page, not just the claim
Claude No hotel-specific figure. Does not browse by default and is calibrated to decline rather than guess Fewer fabrications, more refusals; stale data when web access is off By declining, which is the safer failure Medium: useful for policy text, not for availability
Google Hotels or Booking.com search Not an AI recommendation. Inventory-backed listings Listing quality and review integrity issues, but the property exists Transparently Low: this is the check, not the thing being checked

The row that matters most is the last one. Every AI in this table is a lead generator. The verification still happens somewhere with actual inventory behind it.

A neatly made bed with dark linen beside a white wooden bedside table Photo by Andrew Neel on Unsplash

Why Do AI Models Invent Hotels That Do Not Exist?

Because they were trained to produce a plausible next token, not a verified one. Ask for a boutique hotel near the old town in Ljubljana and the model generates the shape of a hotel recommendation: a name that sounds right for the region, a neighborhood reference, a price signal, a short description. Existence is not a variable in that process. It produces the same fluent output whether the property is real, closed, or has never existed.

Boutique and independent properties are most exposed for two reasons. They have thinner web presence, so the model has less signal and fills the gap from pattern. And their details go stale fastest: seasonal hours, ownership changes, renovation closures and amenity changes all move faster than any training corpus. Rick Steves' team put the consequence plainly in a published caution, warning that AI "might send you to a sight that's closed or have a poor understanding of how much time or distance is required," because it mines the internet and inherits every error in that pool.

Perplexity reduces one dimension of this by citing sources, and introduces another. The researchers who ran the largest published test of AI search attribution describe what they found in these terms:

"Most of the tools we tested presented inaccurate answers with alarming confidence, rarely using qualifying phrases such as 'it appears,' 'it's possible,' 'might,' etc."

Source: Klaudia Jaźwińska and Aisvarya Chandrasekar, researchers at the Tow Center for Digital Journalism, Columbia University, "AI Search Has a Citation Problem," Columbia Journalism Review, March 6, 2025.

That sentence is the mechanism behind the hotel problem in one line. The tone does not vary with the confidence. A hotel the model is certain about and a hotel it assembled from pattern arrive in the same register, with the same fluency, sometimes with the same citation formatting.

What Do the Hallucination Benchmarks Actually Measure?

They measure whether a model contradicts a document you handed it, which is not the question you are asking when you request a hotel. Vectara's leaderboard, the most-cited hallucination benchmark in the industry, scores grounded summarization: the model gets a source document and is scored on whether its summary stays faithful to it. Its own README states the scope directly, saying it is "not evaluating the quality of the summaries, only the factual consistency of them." A hotel recommendation is the opposite task. There is no document. The model is generating from memory.

The gap between those two tasks is large enough to make the benchmark numbers misleading if you carry them across. Grounded summarization scores for the current model families sit in a single-digit to low-teens band, and open-domain factual generation is a much harder problem for the same models. Reading a low grounded-summarization score as "this model rarely invents hotels" is exactly the error this section exists to prevent.

The travel-specific number in circulation has the same problem in a different direction. The claim that roughly 9 in 10 AI-generated itineraries contain at least one factual error traces, through Copyleaks and Generali Travel Insurance and a chain of secondary reports, to a single 2024 audit by the UK marketing agency SEO Travel, which had ChatGPT generate 100 two-day itineraries for ten cities and fact-checked them by hand [SYNTHESIS]. In that audit, 52% of itineraries sent a traveler to a venue outside its opening hours and 24% recommended a permanently closed business. It is one model, one year, standard chat mode, agency research rather than peer-reviewed work. Useful as an order of magnitude, not as a per-model rate.

The honest bottom line: no peer-reviewed study has published a clean hotel-specific hallucination rate for ChatGPT, Gemini or Perplexity. Anyone quoting a precise per-model hotel hallucination percentage without a source is making that number up. And no 2026 replication of any of this work exists, which is itself worth knowing before you weight a 2025 figure heavily.

See how ChatGPT, Gemini and Claude compare on full trip planning tasks

Man standing beside counter Photo by Helena Lopes on Unsplash

When Should I Trust ChatGPT for Hotel Research?

Trust it to shortlist, never to confirm. ChatGPT is genuinely good at narrowing a field: describe the neighborhood feel, the budget, the noise tolerance and the walk you want to the metro, and it will produce a set of candidates that are the right shape. Every name in that set is a hypothesis. Its characteristic hotel failures are inventing plausible boutique names, attaching amenities a property does not have, and quoting price bands that were true at some earlier point.

Turn web browsing on before you use it for accommodation at all. Without browsing it is drawing on training data alone with no live source and no recency, which is the highest-risk configuration for this task. With search enabled, it scored a 67% citation error rate in the Tow Center's news test [SYNTHESIS], so the citations it does produce are not a substitute for checking.

Use it for: shortlisting by vibe and constraint, drafting the questions to ask a property, comparing neighborhoods, summarizing a long cancellation policy you paste in. Do not use it for: confirming a property exists, confirming amenities, current rates, availability, or anything you would pay against without a second source.

When Should I Trust Gemini for Hotel Research?

Trust it where its Google integration is genuinely doing work: current opening hours, a property's presence on Google Maps, recent reviews, whether a neighborhood is walkable to what you want. That live-data path is real and it is Gemini's advantage over a model generating from training data alone.

Be more careful with everything else. The Weldborough case is instructive precisely because it was not a chatbot conversation: an AI content tool published an invented attraction on a real tour operator's website, where it was then indexed, read as authoritative, and acted on by travelers. That is the loop that makes AI hotel errors durable. AI-generated content enters the web, gets cited by the next model, and the error acquires a source.

Gemini's second failure mode is quieter and worth naming: because it favors well-documented properties, a genuinely good small guesthouse with a thin web footprint can be sidelined in favor of a worse but better-indexed option. That is not a hallucination, but it shapes what you see.

Use it for: current hours, map presence, recent reviews, live rates as a first approximation, checking whether a property Google knows about matches what another AI told you. Do not use it for: niche or newly opened properties, star ratings and categories, amenity lists, anything outside major chains without a second check.

When Should I Trust Perplexity for Hotel Research?

Trust it more than the other two, and still open every link. Perplexity is the most defensible starting point because it cites sources at the sentence level, and it had the lowest citation error rate of eight tools in the Tow Center test at 37% [SYNTHESIS]. Its errors are auditable in a way the others' are not: you can see which page it says supports a claim, and go read that page.

Its characteristic failure is the reason you have to. Perplexity will link to a real, live, high-authority URL and attach to it a detail that page does not contain. The hotel name may be right and the rooftop bar invented, with a citation pointing at the hotel's own site, which does not mention a rooftop bar. That is harder to catch than a fabricated name because the citation looks perfect until you read it.

Use it for: first-pass research where you intend to check the work, comparing what different sources say about the same property, tracing a claim back to where it came from. Do not use it for: taking a cited claim at face value. The citation is a pointer to check, not evidence that the check was done.

When Should I Use Google Hotels or Booking.com Instead of Any AI?

Use them the moment you need existence, availability or price to be true rather than plausible. Those platforms are backed by inventory. When a property appears on Booking.com with rooms for your dates, that is a fact about the world rather than a fact about a language model's training distribution. Every AI in this comparison is upstream of that check, and none of them replaces it.

The practical division of labor: ask an AI what kind of place you want and where, then find the specific property through Google Hotels, Booking.com, or the hotel's own site. If you want human judgment rather than a listing, a destination subreddit such as r/JapanTravel or r/PortugalTravel will give you named recommendations from people who stayed there, which is a different and often better signal than either an AI or an aggregate score.

Go straight to these when the AI recommendation returns nothing on your first check. Do not ask the AI for another suggestion in the same conversation; the same process that produced the first name will produce the second.

Black and silver ip desk phone on brown wooden desk Photo by Tomeo Sonner on Unsplash

The Travel Anywhere Hotel Verification Stack for 2026

Three checks, run in this order, before you build any part of a trip around an AI hotel recommendation. In our editorial use they take a few minutes per property, and the first one settles most cases on its own.

Check 1: Two independent booking platforms. Search the exact name on Booking.com and Google Hotels. Legitimate properties appear on at least one. Zero results on both is a strong signal the property does not exist or has closed. Present on one and not the other means look hard at the address, phone number and photo set for discrepancies.

Check 2: Google Maps and street view. Search the name plus the city. You want a pin, a business listing and ideally street-level imagery. Cross-check the address the AI gave you against what Maps shows, because the same-name-wrong-street error is common even for real hotels. If the imagery shows a vacant building or a different business, trust your eyes over the description.

Check 3: The property's own website. Search the name plus "official site." A real hotel almost always has one, with a matching address and contact details. Fabricated properties fail this immediately. For real ones, it is also where you verify the amenities the AI described against what the property itself claims.

If any check comes up empty or shows a mismatch, stop. Do not book on the AI recommendation alone, and do not go back to the same AI for a replacement.

Travel Anywhere addresses this at the workflow level rather than the prompt level: instead of asking a general-purpose model to produce hotel names and hoping they are real, it cross-references suggestions against live booking data before surfacing them.

How Do Real Travelers Catch a Fake Hotel Before Booking in 2026?

They recognize which of four error types they are looking at, because each one fails a different check. The Weldborough case is the dramatic type, and it is also the rarest. The other three cost more travelers more money because they happen to real properties and survive a casual look.

Outright invented properties. The model generates a plausible name that does not exist. Fodor's 2026 reporting on AI travel scams describes travelers arriving at boutique hotels in places like Costa Rica to find empty buildings, with AI-generated reviews having made the fictional property look legitimate. Caught by: Check 1, immediately.

Real name, wrong details. The hotel exists; the price tier, the rooftop pool, the airport shuttle or the pet policy do not. This is the most common and the least visible, because the property passes every existence check and the invented detail only surfaces at the desk. Caught by: Check 3, against the property's own site.

Real name, wrong city. Chains with the same brand in multiple cities are prone to this. An AI can accurately describe a Marriott Courtyard and place it at an address that belongs to a different city's property, or blend two locations into one description. Caught by: Check 2, address against map pin.

Closed or rebranded properties. Training data does not expire, and a hotel that closed in 2023 still exists in forum posts, old reviews and stale listings. It will look convincing and may even appear in older platform records. Caught by: Check 2, street view, plus recent review dates.

For prompts designed to make the model surface these problems itself before you start checking, see prompts to catch AI travel hallucinations. For the itinerary-level version of the same comparison, the deep research itinerary test covers how the three tools handle a full 10-day plan.

FAQ: AI Hotel Hallucinations in 2026

What is the hallucination rate for AI hotel recommendations?

No study has published one. The closest available proxy is citation accuracy on open-ended queries: the Tow Center's March 2025 study found Perplexity at 37% error, ChatGPT Search at 67% and Grok 3 at 94% [SYNTHESIS], testing news attribution. Grounded summarization benchmarks report much lower figures, but a hotel query is not a summarization task, so those numbers understate the risk rather than describing it.

How do I verify that an AI-recommended hotel is real?

Search the exact name on Booking.com and Google Hotels, then look for a matching business pin and street-level imagery on Google Maps, then find the property's own website and confirm the address and contact details match. If any of the three fails or conflicts, do not book on the AI recommendation.

Is Perplexity more reliable than ChatGPT for hotel recommendations?

It is more auditable, which is the more useful property. Its citation error rate was lower in the Tow Center test and it links to sources you can open. But its specific failure mode, attaching fabricated details to real URLs, means you still have to read the cited page rather than trust that a citation exists. Neither is reliable enough to skip verification.

Why does AI invent hotels that do not exist?

Language models predict likely text from training patterns. A hotel query activates the pattern for how hotel descriptions are structured, and the model fills that pattern whether or not the result corresponds to a real property. There is no internal existence check. Obscure and boutique properties with thin web presence are the most likely to be fabricated because the model has the least signal to work from.

Can Gemini recommend fake hotels?

Yes. The best-documented recent case involved an AI content tool on the Tasmania Tours website recommending nonexistent hot springs at Weldborough, reported by ABC News Australia and CNN in January 2026. Google integration reduces but does not eliminate the risk for well-documented properties, and Gemini's more common failure is wrong details on real properties rather than wholly invented ones.

What should I do if a hotel only exists in an AI recommendation?

Treat it as fabricated until proven otherwise, run the three checks, and if none confirms the property, search on a booking platform directly instead of asking the AI again. The same process that generated the first plausible name will generate a second one.

Are AI hotel reviews trustworthy?

No. AI can generate fake reviews with invented guests and stays, and Fodor's 2026 reporting notes that AI-generated reviews have "eliminated classic red flags of fraud: typos and awkward phrasing." Use verified-stay reviews on major platforms, which Booking.com marks explicitly, rather than AI review summaries.

Does AI hallucinate more about boutique hotels than chains?

Yes, consistently, and for a structural reason: less web presence means less training signal, so the model does more filling in. Chains are described from a large, consistent corpus. A twelve-room guesthouse that opened last year may have almost no footprint, which is exactly the condition under which a model produces its most confident inventions.

Bottom Line: The 2026 AI Hotel Recommendation Decision

ChatGPT can and does make up hotels, and so do Gemini and Perplexity, and nobody has measured how often. The number you will see quoted is either from a news-citation study, a grounded-summarization leaderboard, or a 2024 audit of 100 ChatGPT itineraries, and none of those three measures hotel recommendation. That absence is the finding, and it is the reason the verification step is not optional.

Use Perplexity when you intend to check the citations, Gemini when live Google data is genuinely doing the work, ChatGPT with browsing on for shortlisting, and Booking.com or Google Hotels for anything you are about to pay for. Run the three checks per property. Two booking platforms, a map pin, an official site.

Travel Anywhere is working to close this gap at the source by integrating live booking data into the recommendation flow, so existence and availability are part of the answer rather than a separate task left to you.

Ready to make this trip happen? Travel Anywhere plans and books everything, start to finish. Begin at travelanywhere.chat.

Sources

Rachel Caldwell

Rachel CaldwellEditorial Director, TravelAnywhere

Rachel Caldwell is the Editorial Director of TravelAnywhere. She leads the editorial team behind every guide on travelanywhere.blog, focusing on primary research, honest budget math, and recommendations the team would book themselves. Last reviewed August 27, 2026.