AI Travel Index Changelog
Every change to the Tripstitch AI Travel Index, newest first: models tested and the score each entered the field at, suite versions, and changes to how anything is measured.
Re-tested Gemini 3.6 Flash (medium effort) re-tested 89.5
Now at 89.5 on the Travel Score. Recommends a real place 89.1% of the time from memory.
Re-tested Claude Opus 5 (medium effort) re-tested 88.9
Now at 88.9 on the Travel Score. Recommends a real place 84.5% of the time from memory.
Re-tested GPT 5.6 Terra (medium effort) re-tested 87.7
Now at 87.7 on the Travel Score. Recommends a real place 88.8% of the time from memory.
Re-tested GPT 5.6 Luna (medium effort) re-tested 86.4
Now at 86.4 on the Travel Score. Recommends a real place 89.3% of the time from memory.
Re-tested Claude Sonnet 5 (medium effort) re-tested 83.6
Now at 83.6 on the Travel Score. Recommends a real place 71.5% of the time from memory.
Re-tested Grok 4.5 (medium effort) re-tested 83.5
Now at 83.5 on the Travel Score. Recommends a real place 84.5% of the time from memory.
Re-tested Muse Spark 1.2 (medium effort) re-tested 79.1
Now at 79.1 on the Travel Score. Recommends a real place 85.6% of the time from memory.
Re-tested Gemini 3.5 Flash Lite (medium effort) re-tested 75.8
Now at 75.8 on the Travel Score. Recommends a real place 79.2% of the time from memory.
Re-tested Qwen3.8 Max (medium effort) re-tested 75.0
Now at 75.0 on the Travel Score. Recommends a real place 70.1% of the time from memory.
Re-tested Mimo V2.5 Pro (medium effort) re-tested 67.5
Now at 67.5 on the Travel Score. Recommends a real place 77.9% of the time from memory.
Re-tested Deepseek V4 Pro (medium effort) re-tested 66.3
Now at 66.3 on the Travel Score. Recommends a real place 85.7% of the time from memory.
Re-tested Mistral Small 2603 (medium effort) re-tested 63.9
Now at 63.9 on the Travel Score. Recommends a real place 72.5% of the time from memory.
Re-tested GLM 5.2 (medium effort) re-tested 63.8
Now at 63.8 on the Travel Score. Recommends a real place 77.6% of the time from memory.
Suite update Suite 1.6: the search tool stops punishing models for its own crash
The grounded-search task's place-search tool used to crash on a missing or non-finite coordinate, and the whole call was floor-scored for it. It now returns an error the model can read and retry on. Every entrant was re-measured on that task and folded into its published figures; five entrants' grounding had been depressed by the crashes, mistral-small's by over five points.
Launch The AI Travel Index goes live
First public leaderboard, running suite 1.5 against Overture Places, Apple Maps routing and the tz database. No language model judges any answer.
Back to the leaderboard, or read the methodology.