AI Travel Index Changelog

Every change to the Tripstitch AI Travel Index, newest first: models tested and the score each entered the field at, suite versions, and changes to how anything is measured.

  1. New result GPT 5.6 Terra (medium effort) tested 87.5

    Enters the field at 87.5 on the Travel Score. Recommends a real place 88.8% of the time from memory.

  2. New result Gemini 3.6 Flash (medium effort) tested 89.7

    Enters the field at 89.7 on the Travel Score. Recommends a real place 89.1% of the time from memory.

  3. New result Grok 4.5 (medium effort) tested 86.9

    Enters the field at 86.9 on the Travel Score. Recommends a real place 84.5% of the time from memory.

  4. New result GPT 5.6 Luna (medium effort) tested 86.1

    Enters the field at 86.1 on the Travel Score. Recommends a real place 89.3% of the time from memory.

  5. New result Claude Sonnet 5 (medium effort) tested 83.3

    Enters the field at 83.3 on the Travel Score. Recommends a real place 71.5% of the time from memory.

  6. New result Gemini 3.5 Flash Lite (medium effort) tested 75.7

    Enters the field at 75.7 on the Travel Score. Recommends a real place 75.5% of the time from memory.

  7. New result Deepseek V4 Pro (medium effort) tested 67.8

    Enters the field at 67.8 on the Travel Score. Recommends a real place 85.7% of the time from memory.

  8. New result GLM 5.2 (medium effort) tested 65.8

    Enters the field at 65.8 on the Travel Score. Recommends a real place 77.6% of the time from memory.

  9. New result Mistral Small 4 (medium effort) tested 62.0

    Enters the field at 62.0 on the Travel Score. Recommends a real place 72.5% of the time from memory.

  10. Launch The AI Travel Index goes live

    First public leaderboard, running suite 1.5 against Overture Places, Apple Maps routing and the tz database. No language model judges any answer.

Back to the leaderboard, or read the methodology.