AI Guide Reliability Rankings by Roman Ruins Site — An Honest Review
Last Saturday, I was sitting alone on a bench at the Forum Romanum when I asked GPT, "Explain exactly what this column was." I got a fairly convincing answer. The problem is that a graduate student in archaeology standing next to me pointed out the answer was half wrong. That's when it hit me — why does AI make a decent guide at some ruins but talk complete nonsense at others?
It's a bit funny for someone who works in machine learning to be asking this. I know the limitations of the things I build better than anyone. Still, living in Rome, I have a habit of poking around ruins every weekend, and over the past few months I've actually been walking around various sites with an LLM chatbot open. Here's an honest summary of what I found.
The Pantheon — Where AI Is Surprisingly Good
Standing in front of the Pantheon, you're overwhelmed even without a guidebook, but if you ask the AI, "What happens to the floor when rain comes through the oculus?" it explains the drainage system fairly accurately. Because the architectural structure is well-documented and there's a wealth of English-language material, LLMs really shine here. In the three or so tests I ran, the AI covered everything from the changing concrete composition of the dome to the story of the copper being stripped away — on par with a human guide.
Once you get into the details of the tombs inside, though, confident nonsense appears. When I asked it to interpret the Latin inscription on Raphael's tombstone, it produced a translation that sounded plausible but was subtly wrong. A textbook hallucination.

Ostia Antica — Where AI Is Genuinely Helpful
Ostia Antica gets fewer tourists than the ruins in the city centre, making it hard to find guided tours. I love this place. Walking through the ancient shopping streets under the shade of pine trees in summer is a world apart from the crowds at the Forum. With the AI guide running as I walked, explanations like "This was a fuller's laundry" or "This mosaic is a guild mark" were mostly correct. This is where I realised the most practical use case: AI filling in the gaps where on-site information panels are lacking across a vast archaeological site.
That said, the AI will never tell you something like, "If you sit under this tree, around three in the afternoon the light slants through the columns and it's perfect for photos." That kind of thing is the domain of locals.
The Roman Forum — Where AI Fails Most Spectacularly
This is where it falls apart. The layers are just too complex. Republican-era structures with Imperial-era buildings on top and medieval churches nestled over those — when you ask the AI, "What is this brick structure right in front of me?" the answer comes back off by a century or two. Even providing GPS coordinates doesn't help. Once, standing in front of the podium of the Temple of Saturn, I asked a question and got a lengthy explanation of the Temple of Castor and Pollux. Delivered with absolute confidence.
What I learned here is that AI is strong with context-free information but miserably weak when it comes to the physical context of "right here, right now." The moment a human guide points a finger at an arch and says, "See those scratch marks? That's where they quarried the stone in the Middle Ages" — that kind of value simply can't be replaced yet.

The Appian Way — An Unexpected Draw
On a Sunday morning, cycling down the car-free Appian Way with earbuds in, I peppered the AI with questions. It explained the general history of the roadside tombs well. Its account of the catacomb structures wasn't bad either. But when I asked, "Are these wheel ruts in the paving stones really from ancient carts?" the AI said "Yes," while an archaeology guide I met later laughed and said, "Some are, some are from later carts." The correct answer is always "It depends," but AI can't stand that kind of ambiguity.
After months of informal experimenting, a pattern emerges for when an AI guide is actually useful: abundant source material, straightforward site layout, and minimal need for visual context. Places like the Pantheon or Ostia Antica fit perfectly. Conversely, at a site like the Roman Forum — where every stone underfoot is the subject of debate — AI is at its most dangerous, because it gets things wrong without the slightest hesitation.
Next weekend I'm thinking of visiting Villa Adriana. I'm curious how much the AI will hold forth on Emperor Hadrian's architectural tastes. It'll probably be confidently wrong, but that's part of the fun.
Comments