Travel Planning Benchmark: Which AI Models Actually Know Geography
An open benchmark of large language models on travel and location tasks. Every score is checked against Overture Maps, Apple Maps routing and the tz database, with no LLM judging any answer. Hallucination rates, route optimality, schedule feasibility, cost and latency for every model tested.
No results have been published yet. Run php artisan ai:benchmark <model> and drop the file into data/benchmarks/.